Description
Distributed Systems Engineer @ the company
Mission
The company builds persistent computers for AI agents. Our flagship product, Dedalus Machines, gives agents an isolated environment where they can run software, keep files and state, and work over time.
We’re building the persistent compute layer that powers the next generation of autonomous software. Our platform spans distributed storage, virtualization, orchestration, networking, scheduling, and runtime infrastructure for long-running AI agents.
We’re looking for engineers who enjoy designing systems that continue working long after individual machines fail.
You might be a fit if you
Think distributed systems are one of computer science’s most beautiful subjects.
Care deeply about consistency, fault tolerance, and system correctness.
Enjoy designing systems before writing them.
Think latency, throughput, and reliability are all product features.
Read systems papers because they’re genuinely interesting.
Have strong opinions about storage engines, consensus algorithms, scheduling, or distributed architecture.
Believe simple systems are usually harder to build than complicated ones.
Think every abstraction has a cost.
Measure before optimizing, then optimize relentlessly.
View infrastructure as a product for other engineers.
Are high agency and fiercely independent.
Say how things ought to be built, then build them.
Are a competitive teammate with a heart of gold.
Are hungry to learn, improve, and reflect deeply on feedback.
Go above and beyond in everything you do.
What you’ll build
Distributed infrastructure for large-scale AI agent workloads.
Persistent compute and distributed storage systems.
Scheduling and orchestration platforms.
Virtualization and sandboxing infrastructure.
Reliable multi-tenant cloud systems.
Internal developer platforms and infrastructure tooling.
Production systems operating under real-world scale, latency, and fault tolerance constraints.
Representative Projects
You might find yourself working on problems like:
Designing distributed storage systems for persistent agent state.
Building scheduling infrastructure that efficiently allocates compute across thousands of concurrent agents.
Improving reliability and fault tolerance across distributed infrastructure.
Designing virtualization and isolation systems for secure multi-tenant execution.
Optimizing bottlenecks across networking, storage, scheduling, and runtime layers.
Building infrastructure that makes operating AI agents dramatically simpler for developers.