Naive is building autonomous companies: agent systems capable of operating real businesses end to end.
We release autonomous company templates and benchmarks, alongside a studio—where users can deploy and operate these systems. Our infrastructure platform, Vetta, powers long-running, high-volume agent workloads.
Responsibilities
Build real-world agent environments, benchmarks and autonomous company blueprints—and make Vetta the best-performing agent within them.
Turn complex business workflows into reproducible agent environments
Design tasks, rewards, evals and benchmarks that measure real outcomes
Build production-ready autonomous company blueprints
Run experiments, analyze failures and improve agent performance
Develop training data and optimization loops from agent trajectories
Publish credible benchmarks, technical reports and demos
Requirements
Strong Python and research-engineering ability
Experience building agents, evals or RL environments
Deep understanding of tool use, long-horizon tasks and LLM failure modes
Ability to design rigorous experiments and ship production systems
High agency and comfort working on ambiguous 0→1 problems
Nice-to-Haves
RL or post-training experience
Browser, coding or computer-use agent experience
Experience publishing benchmarks or technical research
Familiarity with distributed agent infrastructure
P.S. If you’ve made it to the bottom of this listing and are serious about every point on this role, send Sean a LinkedIn connect with a note.