I'm a Research Scientist with a strong focus on Reinforcement Learning, working at the intersection of theory and application since 2019 at the Center of Excellence in Artificial Intelligence (CEIA) and the Advanced Knowledge Center of Immersive Technologies (AKCIT) in Goiânia, Brazil. I'm also a cofounder of Szero.
My work spans fundamental and applied research — leading R&D projects with private companies across industries such as finance, agribusiness, and tourism to design intelligent systems that deliver real-world impact, taking models from initial data analysis all the way to production.
I'm currently open to PhD opportunities in reinforcement learning and related areas.
My research is rooted in the pursuit of General Artificial Intelligence — enabling agents to learn complex, adaptive behaviors through interaction and experience. I'm particularly drawn to the connection between learned policies and human behavior, exploring how intelligent systems can generalize knowledge, cooperate with humans, and make decisions in dynamic environments.
Sequential decision-making, policy optimization, offline RL, and RL for real-world systems.
Learning-based control, sim-to-real transfer, and embodied agents.
Reasoning, uncertainty, and building language models that behave reliably and safely.
Personalization, fairness, and RL-based recommendation in production environments.
Recommender Systems are widely used to filter information and suggest personalized content to users. Traditionally, these systems focus on maximizing relevance for individual users, which can create popularity bias and lack of diversity in recommendations. This work proposes a recommendation framework, A2Fair, that seeks to balance recommendation relevance with fairness in the exposure of recommended items. The approach consists of an Actor-Critic Reinforcement Learning algorithm, where state representation incorporates features of the last items interacted by the user, as well as the current fairness situation among item groups. The action suggests the next item to recommend, while the adaptive reward function weighs relevance and fairness according to the user's affinity for diversity. The A2Fair framework enables personalized and adaptive recommendations, taking into account individual preferences while promoting fairness and diversity. Experiments conducted on public datasets demonstrated the effectiveness of the A2Fair framework compared to other methods.
In large-scale problems, reinforcement learning systems must use parameterized function approximators, such as neural networks, to generalize across similar situations and actions. Despite proving to be an effective method in specific complex problems, they still fail in generalization, and the most common reinforcement learning evaluation environments still encourage training and evaluation on the same set of environments. To assess an algorithm's ability to generalize across tasks, evaluation environments that measure performance on a distinct test set from those used in training are necessary. This work evaluates the performance of the Proximal Policy Optimization (PPO) algorithm in the General Video Game Artificial Intelligence (GVGAI) evaluation environment, which provides subdivision of a virtual game world into different phases or levels. Although PPO generally reports excellent results, it is notable that the algorithm suffers from overfitting to the training set.
As part of Pequi Mecânico — UFG's competitive robotics team — I collaborated on building autonomous systems for the Latin American Robotics Competition (LARC). Beyond the competitions, I helped bring robotics into the broader community through hands-on Arduino workshops.
Taught technology classes to high school students, introducing them to programming and robotics fundamentals. The curriculum included hands-on projects, such as building Android apps with MIT App Inventor, creating games with Stencyl, and preparing robots for the Brazilian Robotics Olympiad (OBR).
Active participation in R&D initiatives, collaborating with companies to develop innovative AI technologies. From recommendation engines to reinforcement learning systems, these projects bridge academic research with production-grade deployments.
Building high-fidelity digital twins and RL pipelines for humanoids and quadrupeds — covering sim-to-real transfer, motion capture retargeting, offline→online RL, multi-agent coordination, and policy decomposition for edge deployment.
Developing digital agents capable of planning, acting, cooperating, and learning in complex environments — exploring model-based RL, hierarchical planning, multi-agent intention inference, offline policy initialization, and meta-learning for task-conditioned policies.
Implementing RL techniques to orchestrate debtor contact strategies, optimizing the timing, channel, and approach to maximize recovered value on outstanding debts.
Applying reinforcement learning to improve conversion rates on newly negotiated debts, learning optimal follow-up strategies that adapt to individual payment behaviors.
Building recommendation systems to personalize offers across agricultural products — from land and livestock to job listings and specialized services for the agribusiness sector.
Feel free to reach out — whether it's about research, collaboration, or just to say hi.