You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Company hidden

Software Engineer, Infrastructure

  • Office
  • 3-6 years

Salary

Not stated

Similar roles pay $145K - 190K a year · our estimate

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

The company builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at global scale and are used to build millions of real applications.

We're solving the hard part of AI-driven software creation: correctness, reliability, security, and scale in real production systems. The team is built by repeat founders, Olympiad medalists, IIT & IIM alumni, and leaders from Google, Amazon, and Dropbox.

We're hiring builders who want ownership, speed, and impact at global scale.

 

Responsibilities

  • Platform and Infrastructure
  • Maintain stability of our platform consisting of distributed microservices closely interacting with Kubernetes and cloud providers (GCP, AWS)
  • Manage Kubernetes workloads with ArgoCD (GitOps), deploy, monitor, and troubleshoot application syncs, resource trees, and rollouts
  • Debug and resolve complex Kubernetes issues across clusters
  • Manage CDN and edge infrastructure (Cloudflare) for performance, caching, and traffic management
  • Automate infrastructure lifecycle operations and workflows
  • Observability and Incident Response
  • Own the observability stack: Grafana (dashboards, Loki logs, Prometheus metrics) and New Relic (APM, golden metrics, transaction analysis)
  • Enhance monitoring, alerting, and distributed tracing across services
  • Participate in on-call rotation via PagerDuty, handle incident response, and perform root cause analysis
  • Proactively identify reliability risks before they become incidents
  • AI Agent Infrastructure
  • Support the platform that runs AI agent workloads including job scheduling, trajectory tracking, environment provisioning, deployments, and cost attribution
  • Develop Kubernetes controllers and operators to extend platform capabilities for agent orchestration
  • Collaboration and Internal Tooling
  • Work closely with product and backend teams to ensure platform scalability and reliability
  • Build internal tools, automate workflows, and integrate systems to improve team productivity
  • Stay current with Kubernetes releases, CNCF ecosystem updates, and cloud-native best practices

Requirements

  • Core Requirements
  • 3+ years of software/platform engineering experience with production systems
  • Strong proficiency in Go or Python, you write production code in at least one daily
  • Hands-on experience building and deploying services on Kubernetes, not just YAML, you've developed something that runs on K8s
  • Experience with GitOps tooling (ArgoCD, Flux, or similar)
  • Systems Fundamentals
  • Strong networking and DNS fundamentals: TCP/IP, HTTP, load balancing, DNS resolution, TLS, and debugging connectivity issues
  • Solid Linux/OS fundamentals: process management, filesystem, memory, systemd, and comfortable debugging with tools like strace, tcpdump, and netstat
  • Data and Messaging Infrastructure
  • Relational databases: experience with PostgreSQL, MySQL, or similar; indexing, query optimization, replication, and backup/restore procedures
  • NoSQL databases: familiarity with MongoDB, DynamoDB, Redis, or similar for document/key-value workloads
  • Caching: experience with Redis, Memcached, or similar for application and infrastructure-level caching
  • Message queues and streaming: hands-on with Kafka, SQS, RabbitMQ, or similar for event-driven architectures
  • Strong SQL skills for debugging and operational queries
  • Infrastructure and Observability
  • Comfortable with the CNCF ecosystem: Helm, Kustomize, cert-manager, Ingress controllers, CNI/CSI interfaces
  • Hands-on with at least one observability stack (Grafana/Prometheus/Loki, New Relic, Datadog, or similar)
  • Familiarity with GCP and/or AWS: managed Kubernetes (GKE/EKS), networking, IAM, storage, and cloud-native services (SES, SQS, S3, etc.)
  • Experience with CDN/edge platforms (Cloudflare, CloudFront, or similar)
  • Good to Have:
  • Experience building Kubernetes Operators (kubebuilder, operator-sdk, or controller-runtime)
  • Experience tuning Kubernetes core components (API server, kubelet, scheduler)
  • Familiarity with AI/LLM infrastructure: token management, cost tracking, agent orchestration
  • Experience with CI/CD pipelines (GitHub Actions, automated testing, deployment pipelines)
  • Infrastructure as Code experience (Terraform, Pulumi, or similar)
  • Previous work on large-scale distributed systems or platform-as-a-service
  • Startup experience, you thrive in fast-paced, ambiguous environments
  • A generalist who can context-switch between debugging a K8s deployment, setting up a Grafana alert, and configuring CDN rules, all in the same day
  • You enjoy solving complex infrastructure challenges and automating away toil
  • You dig deep, when something breaks, you find the root cause, not just the workaround
  • You communicate clearly and can collaborate effectively in a fast-moving, distributed team
  • Tech Stack: We don't require previous experience with our entire stack, but enthusiasm for learning is key: Go, Python, Kubernetes, ArgoCD, Helm, GCP, AWS, Cloudflare, Grafana, Prometheus, Loki, New Relic, PagerDuty, PostgreSQL, MongoDB, Redis, Kafka, and GitHub.

Benefits

  • 401(k)
  • Health, dental, and vision insurance
  • Unlimited Paid Time Off: take the time you need to recharge and come back refreshed
  • Flexible Working Hours: work arrangements that fit your life and commitments
  • Let's build the future of software together.
  •  
  •  
  • Experience: Any (new grads ok)
  • Visa: US citizen/visa only

Where you’d work

From the office

The office

No visa sponsorship

About the company

Company hidden

  • Industry: AI

Office in San Francisco, United States

Your chances

Likely still hiring, not crowded yet, and a person reads your message.

  • 17 checks run
  • 5 good signs
  • 2 red flags

Still hiring?

12 checks

Likely active

In its favour2

  • Still on the company's own careers site, checked 4 h agoModerate evidence
  • A hiring contact is attached to itSlight evidence

Against it1

  • None of the company's 4 open roles was posted in the last 2 weeksModerate evidence

How crowded?

5 checks

Low

In its favour3

  • You can message the hiring contact and skip the queueModerate evidence
  • In the office in San Francisco: only people nearby can take itSlight evidence
  • Asks for Helm, which fewer than 1% of open roles doSlight evidence

Against it1

  • Open for 2 weeks: applications have had time to pile upModerate evidence

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details11 facts · Role, Location, Compensation, Employment, Company
Tech stack
  • Python
  • Go
  • AWS
  • Kubernetes
  • Terraform
  • SQL
  • PostgreSQL
  • Kafka
  • Redis
  • MySQL
  • MongoDB
  • Linux
  • CI/CD
  • LLMs
  • GitHub Actions
  • Argo CD
  • Helm
  • Prometheus
  • Grafana
  • Datadog
Type
Full-time
Industry
AI
Region
United States
Pay period
Annual
Show 6 more factsShow less

Role

Category
DevOps & Infrastructure
Experience
3+ years
Tech stack
  • Python
  • Go
  • AWS
  • Kubernetes
  • Terraform
  • SQL
  • PostgreSQL
  • Kafka
  • Redis
  • MySQL
  • MongoDB
  • Linux
  • CI/CD
  • LLMs
  • GitHub Actions
  • Argo CD
  • Helm
  • Prometheus
  • Grafana
  • Datadog

Location

Work model
Office
Region
United States
Office
  • San Francisco, United States
Visa sponsorship
Not sponsored

Compensation

Salary
Salary by agreement
Pay period
Annual

Employment

Type
Full-time

Company

Industry
AI

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Lead IT Engineer

    Salary by agreement

    • Remote
    • Lead & Manager
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Security Engineer, New Grad

    Salary by agreement

    • Office · Dublin
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Site Reliability Engineer

    Salary by agreement

    • Remote
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Security Operations Analyst

    Salary by agreement

    • Hybrid · Wellington, Auckland
    • Senior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason