Software Engineer, Infrastructure
cognition · Remote
## We are an applied AI lab building end-to-end software agents.
We're the makers of Devin, the first AI software engineer.
Our team is extremely talent-dense. Among our founding team, we have world-class competitive programmers, former founders, and leaders from companies at the cutting edge of AI including Scale AI, Palantir, Cursor, Waymo, Tesla, Lunchclub, Modal, Google DeepMind, and Nuro.
Building Devin is just the first step—our hardest challenges still lie ahead. If you’re excited to solve some of the world’s biggest problems and build AI that can reason on real-world tasks, apply to join us.
## Who We Are
Cognition is an applied AI lab building end-to-end software agents. We are behind Devin, the first AI software engineer, and Windsurf, an AI-native IDE. Our vision is AI that works alongside engineers as a genuine teammate, not a tool.
We are a small, talent-dense team of competitive programmers, former founders, and researchers from Scale AI, Palantir, Cursor, Google DeepMind, and others.
## Role Mission
Infrastructure Engineers at Cognition build the systems that everything else runs on. Devin executes long-horizon tasks across sandboxed environments, spawns subagents, uses tools, and does all of this at scale across millions of sessions. Windsurf serves developers in real time with low-latency editor intelligence and agent-in-the-loop workflows. None of that works without world-class infrastructure underneath it. You will own the compute, orchestration, networking, and platform systems that make both products possible, and you will build them with the same engineering rigor and product instinct that the rest of the team brings to everything else. This is a role for engineers who care deeply about reliability, scale, and the craft of building systems that other engineers love to build on top of.
## What You'll Accomplish
• Own agent execution infrastructure: Design and operate the sandboxed compute environments that power Devin's task execution, including VM orchestration, container management, resource scheduling, and isolation at scale.
• Build and maintain the developer platform: Create the internal infrastructure that Cognition engineers depend on every day: CI/CD pipelines, deployment systems, developer tooling, and the abstractions that let product teams ship fast.
• Drive reliability and observability: Define SLOs, build monitoring and alerting systems, lead incident response, and close the loop on postmortems so the same failure never happens twice.
• Scale with the product: Anticipate capacity and architecture needs before they become bottlenecks; make deliberate infrastructure investments that stay ahead of product growth.
• Partner across engineering: Work closely with product engineers and researchers to understand what the systems they are building require, and design infrastructure that meets those needs without creating unnecessary complexity.
## Exceptional Candidates Have Demonstrated
• Deep systems engineering: Experience designing and operating large-scale distributed infrastructure; strong opinions about reliability, failure modes, and operational excellence.
• Cloud and container expertise: Hands-on proficiency with Kubernetes, cloud platforms (AWS, GCP, or Azure), and infrastructure-as-code tools like Terraform.
• Strong software engineering fundamentals: Infrastructure at Cognition means writing real code; proficiency in Python and comfort owning complex systems codebases.
• Observability instincts: Knows how to instrument systems, build meaningful dashboards, and design alerts that surface signal without noise.
• Security and isolation mindset: Experience with sandboxing, network isolation, and secure multi-tenant compute, particularly relevant for agentic workloads running arbitrary code.
• Relevant industry experience: Prior experience at a frontier AI lab, applied AI company, or developer tools company; you know what good looks like in this category.
• Degree from a top-tier university: BS, MS, or equivalent in Computer Science, Mathematics, Engineering, or a related technical discipline from a highly selective program.
## Compensation & Benefits
• Base Salary:
60,000 -
00,000 + significant early-stage equityWe're the makers of Devin, the first AI software engineer.
Our team is extremely talent-dense. Among our founding team, we have world-class competitive programmers, former founders, and leaders from companies at the cutting edge of AI including Scale AI, Palantir, Cursor, Waymo, Tesla, Lunchclub, Modal, Google DeepMind, and Nuro.
Building Devin is just the first step—our hardest challenges still lie ahead. If you’re excited to solve some of the world’s biggest problems and build AI that can reason on real-world tasks, apply to join us.
## Who We Are
Cognition is an applied AI lab building end-to-end software agents. We are behind Devin, the first AI software engineer, and Windsurf, an AI-native IDE. Our vision is AI that works alongside engineers as a genuine teammate, not a tool.
We are a small, talent-dense team of competitive programmers, former founders, and researchers from Scale AI, Palantir, Cursor, Google DeepMind, and others.
## Role Mission
Infrastructure Engineers at Cognition build the systems that everything else runs on. Devin executes long-horizon tasks across sandboxed environments, spawns subagents, uses tools, and does all of this at scale across millions of sessions. Windsurf serves developers in real time with low-latency editor intelligence and agent-in-the-loop workflows. None of that works without world-class infrastructure underneath it. You will own the compute, orchestration, networking, and platform systems that make both products possible, and you will build them with the same engineering rigor and product instinct that the rest of the team brings to everything else. This is a role for engineers who care deeply about reliability, scale, and the craft of building systems that other engineers love to build on top of.
## What You'll Accomplish
• Own agent execution infrastructure: Design and operate the sandboxed compute environments that power Devin's task execution, including VM orchestration, container management, resource scheduling, and isolation at scale.
• Build and maintain the developer platform: Create the internal infrastructure that Cognition engineers depend on every day: CI/CD pipelines, deployment systems, developer tooling, and the abstractions that let product teams ship fast.
• Drive reliability and observability: Define SLOs, build monitoring and alerting systems, lead incident response, and close the loop on postmortems so the same failure never happens twice.
• Scale with the product: Anticipate capacity and architecture needs before they become bottlenecks; make deliberate infrastructure investments that stay ahead of product growth.
• Partner across engineering: Work closely with product engineers and researchers to understand what the systems they are building require, and design infrastructure that meets those needs without creating unnecessary complexity.
## Exceptional Candidates Have Demonstrated
• Deep systems engineering: Experience designing and operating large-scale distributed infrastructure; strong opinions about reliability, failure modes, and operational excellence.
• Cloud and container expertise: Hands-on proficiency with Kubernetes, cloud platforms (AWS, GCP, or Azure), and infrastructure-as-code tools like Terraform.
• Strong software engineering fundamentals: Infrastructure at Cognition means writing real code; proficiency in Python and comfort owning complex systems codebases.
• Observability instincts: Knows how to instrument systems, build meaningful dashboards, and design alerts that surface signal without noise.
• Security and isolation mindset: Experience with sandboxing, network isolation, and secure multi-tenant compute, particularly relevant for agentic workloads running arbitrary code.
• Relevant industry experience: Prior experience at a frontier AI lab, applied AI company, or developer tools company; you know what good looks like in this category.
• Degree from a top-tier university: BS, MS, or equivalent in Computer Science, Mathematics, Engineering, or a related technical discipline from a highly selective program.
## Compensation & Benefits
• Base Salary:
Careeroza — One-stop Zone for Aspirants
Study material, Careeroza mentorship, tech jobs, and career guidance on careeroza.com.
Public study materials
- Express.js — Web APIs & middleware (expressjs)
- Application setup · basic
- Routing deep dive · basic
- 1. MVC / Layered Architecture · medium
- Validation · medium
- File uploads · medium
- Sessions & auth (stateful) · medium
- Passport & strategies · medium
- Templating & SSR · medium
- WebSockets & SSE · medium
- Security middleware · advance
- Reverse proxies & trust · advance
- Performance · advance
- API design & versioning · advance
- Testing with Supertest · advance
- GraphQL & tRPC (overview) · advance
- Deployment checklist · advance
- Middlewares · basic
- Request & Response · basic
- AWS Crash Course (AWS)
- What is Cloud ? · basic
- What is AWS ? · basic
- If not cloud ? · basic
- Cloud Computing · basic
- AWS Pricing · basic
- AWS Shared Responsibility Model · basic
- AWS Management Console · basic
- AWS SDKs · basic
- AWS IAM · medium
- Users, Groups, Roles · medium
- Policies · medium
- AWS Organizations · medium
- AWS Cognito · medium
- AWS Directory Service · medium
- AWS KMS (Key Management Service) · medium
- AWS Secrets Manager · medium
- AWS Shield · medium
- AWS WAF · medium
- AWS Inspector · medium
- AWS GuardDuty · medium
- EC2 · advance
- Launching EC2 Instances · advance
- EBS Volumes · advance
- Security Groups · advance
- Key Pairs · advance
- Elastic IP · advance
- User Data Scripts · advance
- Auto Scaling · advance
- Load Balancers · advance
- ALB · advance
- NLB · advance
- Serverless Compute ,AWS Lambda, Lambda Layers · advance
- Event-Driven Architecture · advance
- ECS · advance
- EKS · advance
- Django (Django)
- What is Django · basic
- Installing Django · basic
- Features of Django · basic
- MVT Architecture · basic
- Django vs Flask · basic
- Creating Project & Creating App · basic
- Django Project Structure · basic
- URL Routing · basic
- Views · basic
- Templates · basic
- Static & Media Files · medium
- Models · medium
- ORM (Object Relational Mapping) · medium
- Model Relationships · medium
- Migrations · medium
- Django Admin · medium
- Forms · medium
- Authentication · medium
- Authorization · medium
- Middleware · medium
- Signals · medium
- Class Based Views Deep Dive · advance
- Generic Views · advance
- File Handling · advance
- Django REST Framework (DRF) · advance
- Advanced ORM · advance
- Caching · advance
- Asynchronous Django · advance
- Background Tasks · advance
- Interview Questions · interview-questions
- System Designing (System Designing)
- Day-1 : What is system Designing ? · basic
- Day-2 : Vertical vs. Horizontal Scaling · basic
- Day-3:How to do vertical scaling ? · basic
- Day4:How to do horizaontal scaling ? · basic
- Day:5TCP vs UDP · basic
- Day6:IP & DNS · basic
- Day7:Client-Server Model · basic
- Day8:HTTP & HTTPS · basic
- Databases (SQL vs NoSQL) · medium
- Caching · medium
- Day9:Latency & Throughput · basic
- Load Balancing · medium
- Indexes & Query Optimization · medium
- CDN · medium
- Proxies · medium
- Message Queues · medium
- Horizontal vs Vertical Scaling · medium
- Database Replication · advance
- Database Sharding · advance
- Consistent Hashing · advance
- CAP Theorem · advance
- Rate Limiting · advance
- Service Discovery · advance
- Event-Driven Architecture · advance
- API Gateway · advance
- Distributed Consensus · expert
- Microservices · expert
- Observability · expert
- Idempotency · expert
- PACELC Theorem · expert
- Two-Phase Commit · expert
- Back-of-Envelope Estimation · expert
- Designing for Failure · expert
- Node.js — Server-side JavaScript (nodejs)
- Getting started · basic
- JavaScript on the server · basic
- CommonJS modules · basic
- ES modules (ESM) · basic
- npm & package management · basic
- Asynchronous JavaScript in Node · basic
- The event loop · basic
- Essential core utilities · basic
- process & configuration · basic
- File system basics · basic
- HTTP & HTTPS servers · medium
- Streams · medium
- Events & EventEmitter · medium
- Advanced filesystem · medium
- crypto · medium
- Compression & encoding · medium
- Child processes · medium
- net, dgram & DNS · medium
- readline, timers & scheduling · medium
- Testing & diagnostics (intro) · medium
- Worker threads · advance
- cluster & multi-process scaling · advance
- Performance & tuning · advance
- Debugging & observability · advance
- Security hardening · advance
- Native addons & N-API · advance
- Architecture patterns · advance
- Graceful shutdown · advance
- 100 Questions · interview questions
- Python (Python)
- Python Fundamentals · basic
- Control Flow · basic
- Strings · basic
- Collections / Data Structures · basic
- Functions · basic
- Modules and Packages · basic
- File Handling · basic
- Exception Handling · basic
- Object-Oriented Programming (OOP · medium
- Advanced Python Concepts · advance
- Functional Programming · advance
- Multithreading & Multiprocessing · advance
- Async Programming · advance
- JavaScript (JavaScript)
- JS Introduction · basic
- Variables & Data Types · basic
- Operators · basic
- Control Flow · basic
- Functions · basic
- Scope & Execution · basic
- Closures · basic
- Objects · basic
- Arrays · basic
- Strings · basic
- DOM Manipulation · basic
- Browser APIs · medium
- Asynchronous JavaScript · medium
- Fetch & APIs · medium
- ES6+ Features · medium
- OOP in JavaScript · medium
- Prototype & Inheritance · advance
- Advanced Functions · advance
- Memory Management · advance
- Error Handling · advance
- Modules · advance
- Advanced Async Concepts · advance
- Functional Programming · advance
- JavaScript Internals · advance
- Performance Optimization · advance
- MongoDB — Documents & data modeling (MongoDB)
- Introduction · basic
- Shell, Compass & tools · basic
- Databases & collections · basic
- CRUD operations · basic
- Indexes deep dive · medium
- Explain plans & performance · medium
- Aggregation framework · medium
- Schema design patterns · medium
- Mongoose basics · medium
- Mongoose advanced · medium
- Drivers & connection · medium
- Operators for updates & arrays · medium
- Replication & read preferences · advance
- Write concern & read concern · advance
- Multi-document transactions · advance
- Change streams · advance
- Sharding (overview) · advance
- Atlas Search & full-text · advance
- GridFS & large files · advance
- Backup, restore & ops · advance
- SQL (SQL)
- SQL Fundamentals · basic
- Database Operations · basic
- Table Operations · basic
- CRUD Operations · basic
- Filtering & Operators · basic
- SQL Functions · basic
- GROUPING Data · basic
- Joins · medium
- Constraints · basic
- Subqueries · medium
- Set Operators · medium
- Views · medium
- Indexes · medium
- Normalization · advance
- Transactions · advance
- Stored Procedures & Functions · advance
- Triggers · advance
- Advanced SQL · advance
- Query Optimization · advance
- Database Design · advance
- SQL Security · advance
- Backup & Recovery · advance
- Questions · interview questions
- System Architecture (System Architecture)
- Fundamentals of System Architecture · basic
- Distributed System Basics · basic
- System Reliability Concepts · basic
- Scaling Concepts · basic
- Networking Basics · basic
- Web Communication · basic
- API Communication · basic
- Proxy & Delivery Systems · basic
- Web Architecture Basics · basic
- Rendering Architectures · basic
- Frontend Advanced Concepts · basic
- Message Queue Basics · medium
- What is Load Balancer · medium
- Load Balancing Algorithms · medium
- API Design Basics · medium
- API Protection · medium
- Authentication Basics · medium
- Security Tokens · medium
- Security Threats · medium
- Encryption & Security · medium
- SQL Database Basics · medium
- SQL Scaling Concepts · medium
- NoSQL Databases · medium
- Database Optimization · medium
- Replication Strategies · medium
- Caching Basics · medium
- Cache Storage Systems · medium
- Cache Strategies · medium
- Event-Driven Systems · advance
- Queue Reliability · advance
- Microservices Basics · advance
- Microservice Communication · advance
- Distributed Transactions · advance
- DevOps Basics · advance
- Automation Tools · advance
- Deployment Strategies · advance
- Monitoring Basics · advance
- Monitoring Tools · advance
- Distributed System Concepts · advance
- Distributed Algorithms · advance
- Angular (Angular)
- Angular Fundamentals · basic
- Project Structure · basic
- Components & Templates · basic
- Data Binding · basic
- Directives · basic
- Pipes · basic
- Component Communication · basic
- Lifecycle Hooks · basic
- Routing Basics · basic
- Routing · basic
- API Calls · basic
- Forms · medium
- Routing · medium
- Services & Dependency Injection · medium
- RxJS & Observables · medium
- Authentication & Security · medium
- Component Interaction · medium
- State Management Basics · medium
- Error Handiling · medium
- Perfomance Basic · medium
- Real World Features · medium
- Advanced Angular Architecture · advance
- Change Detection · advance
- Advanced RxJS · advance
- State Management · advance
- Dynamic Rendering · advance
- Perfomance Optimization · advance
- Modern Angular · advance
- STAR (Situation, Task, Action, and Result) (Situation Based Questions)
• Medical, Dental, Vision: Fully paid for you and your dependents
• 401(k): Company match included
Perks: Private chef, cozy slippers, endless snacks, and more
## Equal Opportunity
Cognition is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law. We are committed to providing reasonable accommodations for candidates with disabilities throughout the hiring process - please let us know if you need any.