Principal Software Engineer
astronomer · Remote
Experience: 10+ years
Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications. Trusted by more than 800 of the world's leading enterprises, Astronomer lets businesses do more with their data. To learn more, visit www.astronomer.io .
About the Role
We are seeking an exceptional Principal Software Engineer to provide technical leadership, architectural vision, and team mentorship for our core engineering group. In this role, you will be the primary architect and public face of a cloud-agnostic software platform that helps enterprise customers provision, scale, and orchestrate Apache Airflow instances completely on-premises within their own environments.
This is a unique infrastructure challenge: our platform is designed to be completely self-contained, operating without reliance on an external SaaS backbone or mandatory outbound communication. Both our control plane and data planes deploy locally inside the customer’s infrastructure. You will lead the engineering team in defining, designing, and building these systems to function reliably across a spectrum of enterprise network setups—ranging from standard customer-managed clouds to strictly isolated, air-gapped environments.
What You'll Do
• Define, Design, and Build Core Architecture : Own the end-to-end definition, design, and implementation of both the control plane and data plane services, creating highly resilient inter-connections between them to power this self-contained infrastructure platform.
• Architect an Extensible Platform API : Drive the design, data modeling, versioning, and long-term backward compatibility of robust APIs (such as REST or gRPC). These APIs serve as the core platform engine that power the customer-facing UI, and the developer endpoints for customer-facing programmatic integrations.
• Design Comprehensive Platform Observability : Architect the telemetry framework (metrics, logs, and traces) for the self-contained platform itself, ensuring deep visibility into both the core control plane services and the performance, health, and status of the underlying Airflow instances it provisions.
• Architect for Varied Isolation Levels : Oversee the architecture of a self-contained software platform, ensuring all services, images, dependencies, and configurations can run smoothly whether deployed in connected private clouds or highly secure, zero-egress networks.
• Provide Technical Leadership : Act as the technical face of the team, translating complex enterprise infrastructure needs into execution roadmaps while aligning closely with product management and executive stakeholders.
• Mentor and Elevate : Foster a culture of engineering and operational excellence through comprehensive design reviews, high-leverage code contributions, and direct career mentorship for senior engineers.
What You'll Bring
• Extensive Infrastructure Experience : 10+ years of professional software engineering experience, with a significant track record operating at a Staff, Senior Staff, or Principal level within an enterprise infrastructure, database-as-a-service (DBaaS), or private cloud organization.
• Control Plane & API Productization : Proven success building production-grade infrastructure control planes, with a deep understanding of API design principles required to support programmatic customer workflows and user interfaces.
• Cloud-Native Observability Expertise : Practical experience designing and implementing infrastructure observability stacks (e.g., Prometheus, OpenTelemetry, Grafana) to capture, aggregate, and surface system health and workflow data within isolated environments.
• Strong Kubernetes Foundation : Solid operational knowledge and architectural understanding of container orchestration platforms, specifically Kubernetes . You have practical experience building custom Kubernetes Operators (CRDs) , managing Helm deployments, and working with the Kubernetes API to orchestrate application lifecycles.
• Hybrid Cloud Architecture : Hands-on systems design expertise involving isolated environments across multiple public cloud providers (AWS, GCP, Azure), Red Hat OpenShift enterprise deployments, and bare-metal on-premise infrastructure.
• Advanced Coding & Systems Design : Expert-level software development skills (e.g., Go, Python, TypeScript) with deep knowledge of distributed state stores (such as Etcd, Consul, or PostgreSQL), concurrency, and network security patterns.
• Top-Notch Technical Communication : Exceptional oral and written communication skills with a proven ability to author crisp, high-density technical "one-pagers" and actionable RFCs. You can drive fast organizational alignment by distilling complex systems choices into brief, highly readable documents, presenting confidently to executive leadership, and engaging deeply with strategic enterprise clients to unearth their core infrastructure requirements.
The estimated salary for this role ranges from
77,000 -
90,000 based on leveling and geography, along with an equity component and a comprehensive benefits package. This range is merely an estimate; actual compensation may deviate from this range based on skills, experience, and qualifications.About the Role
We are seeking an exceptional Principal Software Engineer to provide technical leadership, architectural vision, and team mentorship for our core engineering group. In this role, you will be the primary architect and public face of a cloud-agnostic software platform that helps enterprise customers provision, scale, and orchestrate Apache Airflow instances completely on-premises within their own environments.
This is a unique infrastructure challenge: our platform is designed to be completely self-contained, operating without reliance on an external SaaS backbone or mandatory outbound communication. Both our control plane and data planes deploy locally inside the customer’s infrastructure. You will lead the engineering team in defining, designing, and building these systems to function reliably across a spectrum of enterprise network setups—ranging from standard customer-managed clouds to strictly isolated, air-gapped environments.
What You'll Do
• Define, Design, and Build Core Architecture : Own the end-to-end definition, design, and implementation of both the control plane and data plane services, creating highly resilient inter-connections between them to power this self-contained infrastructure platform.
• Architect an Extensible Platform API : Drive the design, data modeling, versioning, and long-term backward compatibility of robust APIs (such as REST or gRPC). These APIs serve as the core platform engine that power the customer-facing UI, and the developer endpoints for customer-facing programmatic integrations.
• Design Comprehensive Platform Observability : Architect the telemetry framework (metrics, logs, and traces) for the self-contained platform itself, ensuring deep visibility into both the core control plane services and the performance, health, and status of the underlying Airflow instances it provisions.
• Architect for Varied Isolation Levels : Oversee the architecture of a self-contained software platform, ensuring all services, images, dependencies, and configurations can run smoothly whether deployed in connected private clouds or highly secure, zero-egress networks.
• Provide Technical Leadership : Act as the technical face of the team, translating complex enterprise infrastructure needs into execution roadmaps while aligning closely with product management and executive stakeholders.
• Mentor and Elevate : Foster a culture of engineering and operational excellence through comprehensive design reviews, high-leverage code contributions, and direct career mentorship for senior engineers.
What You'll Bring
• Extensive Infrastructure Experience : 10+ years of professional software engineering experience, with a significant track record operating at a Staff, Senior Staff, or Principal level within an enterprise infrastructure, database-as-a-service (DBaaS), or private cloud organization.
• Control Plane & API Productization : Proven success building production-grade infrastructure control planes, with a deep understanding of API design principles required to support programmatic customer workflows and user interfaces.
• Cloud-Native Observability Expertise : Practical experience designing and implementing infrastructure observability stacks (e.g., Prometheus, OpenTelemetry, Grafana) to capture, aggregate, and surface system health and workflow data within isolated environments.
• Strong Kubernetes Foundation : Solid operational knowledge and architectural understanding of container orchestration platforms, specifically Kubernetes . You have practical experience building custom Kubernetes Operators (CRDs) , managing Helm deployments, and working with the Kubernetes API to orchestrate application lifecycles.
• Hybrid Cloud Architecture : Hands-on systems design expertise involving isolated environments across multiple public cloud providers (AWS, GCP, Azure), Red Hat OpenShift enterprise deployments, and bare-metal on-premise infrastructure.
• Advanced Coding & Systems Design : Expert-level software development skills (e.g., Go, Python, TypeScript) with deep knowledge of distributed state stores (such as Etcd, Consul, or PostgreSQL), concurrency, and network security patterns.
• Top-Notch Technical Communication : Exceptional oral and written communication skills with a proven ability to author crisp, high-density technical "one-pagers" and actionable RFCs. You can drive fast organizational alignment by distilling complex systems choices into brief, highly readable documents, presenting confidently to executive leadership, and engaging deeply with strategic enterprise clients to unearth their core infrastructure requirements.
The estimated salary for this role ranges from
Careeroza — One-stop Zone for Aspirants
Study material, Careeroza mentorship, tech jobs, and career guidance on careeroza.com.
Public study materials
- Django (Django)
- What is Django · basic
- Installing Django · basic
- Features of Django · basic
- MVT Architecture · basic
- Django vs Flask · basic
- Creating Project & Creating App · basic
- Django Project Structure · basic
- URL Routing · basic
- Views · basic
- Templates · basic
- Static & Media Files · medium
- Models · medium
- ORM (Object Relational Mapping) · medium
- Model Relationships · medium
- Migrations · medium
- Django Admin · medium
- Forms · medium
- Authentication · medium
- Authorization · medium
- Middleware · medium
- Signals · medium
- Class Based Views Deep Dive · advance
- Generic Views · advance
- File Handling · advance
- Django REST Framework (DRF) · advance
- Advanced ORM · advance
- Caching · advance
- Asynchronous Django · advance
- Background Tasks · advance
- Interview Questions · interview-questions
- Python (Python)
- Python Fundamentals · basic
- Control Flow · basic
- Strings · basic
- Collections / Data Structures · basic
- Functions · basic
- Modules and Packages · basic
- File Handling · basic
- Exception Handling · basic
- Object-Oriented Programming (OOP · medium
- Advanced Python Concepts · advance
- Functional Programming · advance
- Multithreading & Multiprocessing · advance
- Async Programming · advance
- System Architecture (System Architecture)
- Fundamentals of System Architecture · basic
- Distributed System Basics · basic
- System Reliability Concepts · basic
- Scaling Concepts · basic
- Networking Basics · basic
- Web Communication · basic
- API Communication · basic
- Proxy & Delivery Systems · basic
- Web Architecture Basics · basic
- Rendering Architectures · basic
- Frontend Advanced Concepts · basic
- Message Queue Basics · medium
- What is Load Balancer · medium
- Load Balancing Algorithms · medium
- API Design Basics · medium
- API Protection · medium
- Authentication Basics · medium
- Security Tokens · medium
- Security Threats · medium
- Encryption & Security · medium
- SQL Database Basics · medium
- SQL Scaling Concepts · medium
- NoSQL Databases · medium
- Database Optimization · medium
- Replication Strategies · medium
- Caching Basics · medium
- Cache Storage Systems · medium
- Cache Strategies · medium
- Event-Driven Systems · advance
- Queue Reliability · advance
- Microservices Basics · advance
- Microservice Communication · advance
- Distributed Transactions · advance
- DevOps Basics · advance
- Automation Tools · advance
- Deployment Strategies · advance
- Monitoring Basics · advance
- Monitoring Tools · advance
- Distributed System Concepts · advance
- Distributed Algorithms · advance
- Express.js — Web APIs & middleware (expressjs)
- Application setup · basic
- Routing deep dive · basic
- 1. MVC / Layered Architecture · medium
- Validation · medium
- File uploads · medium
- Sessions & auth (stateful) · medium
- Passport & strategies · medium
- Templating & SSR · medium
- WebSockets & SSE · medium
- Security middleware · advance
- Reverse proxies & trust · advance
- Performance · advance
- API design & versioning · advance
- Testing with Supertest · advance
- GraphQL & tRPC (overview) · advance
- Deployment checklist · advance
- Middlewares · basic
- Request & Response · basic
- System Designing (System Designing)
- Day-1 : What is system Designing ? · basic
- Day-2 : Vertical vs. Horizontal Scaling · basic
- Day-3:How to do vertical scaling ? · basic
- Day4:How to do horizaontal scaling ? · basic
- Day:5TCP vs UDP · basic
- Day6:IP & DNS · basic
- Day7:Client-Server Model · basic
- Day8:HTTP & HTTPS · basic
- Databases (SQL vs NoSQL) · medium
- Caching · medium
- Day9:Latency & Throughput · basic
- Load Balancing · medium
- Indexes & Query Optimization · medium
- CDN · medium
- Proxies · medium
- Message Queues · medium
- Horizontal vs Vertical Scaling · medium
- Database Replication · advance
- Database Sharding · advance
- Consistent Hashing · advance
- CAP Theorem · advance
- Rate Limiting · advance
- Service Discovery · advance
- Event-Driven Architecture · advance
- API Gateway · advance
- Distributed Consensus · expert
- Microservices · expert
- Observability · expert
- Idempotency · expert
- PACELC Theorem · expert
- Two-Phase Commit · expert
- Back-of-Envelope Estimation · expert
- Designing for Failure · expert
- SQL (SQL)
- SQL Fundamentals · basic
- Database Operations · basic
- Table Operations · basic
- CRUD Operations · basic
- Filtering & Operators · basic
- SQL Functions · basic
- GROUPING Data · basic
- Joins · medium
- Constraints · basic
- Subqueries · medium
- Set Operators · medium
- Views · medium
- Indexes · medium
- Normalization · advance
- Transactions · advance
- Stored Procedures & Functions · advance
- Triggers · advance
- Advanced SQL · advance
- Query Optimization · advance
- Database Design · advance
- SQL Security · advance
- Backup & Recovery · advance
- Questions · interview questions
- JavaScript (JavaScript)
- JS Introduction · basic
- Variables & Data Types · basic
- Operators · basic
- Control Flow · basic
- Functions · basic
- Scope & Execution · basic
- Closures · basic
- Objects · basic
- Arrays · basic
- Strings · basic
- DOM Manipulation · basic
- Browser APIs · medium
- Asynchronous JavaScript · medium
- Fetch & APIs · medium
- ES6+ Features · medium
- OOP in JavaScript · medium
- Prototype & Inheritance · advance
- Advanced Functions · advance
- Memory Management · advance
- Error Handling · advance
- Modules · advance
- Advanced Async Concepts · advance
- Functional Programming · advance
- JavaScript Internals · advance
- Performance Optimization · advance
- Angular (Angular)
- Angular Fundamentals · basic
- Project Structure · basic
- Components & Templates · basic
- Data Binding · basic
- Directives · basic
- Pipes · basic
- Component Communication · basic
- Lifecycle Hooks · basic
- Routing Basics · basic
- Routing · basic
- API Calls · basic
- Forms · medium
- Routing · medium
- Services & Dependency Injection · medium
- RxJS & Observables · medium
- Authentication & Security · medium
- Component Interaction · medium
- State Management Basics · medium
- Error Handiling · medium
- Perfomance Basic · medium
- Real World Features · medium
- Advanced Angular Architecture · advance
- Change Detection · advance
- Advanced RxJS · advance
- State Management · advance
- Dynamic Rendering · advance
- Perfomance Optimization · advance
- Modern Angular · advance
- STAR (Situation, Task, Action, and Result) (Situation Based Questions)
- Node.js — Server-side JavaScript (nodejs)
- Getting started · basic
- JavaScript on the server · basic
- CommonJS modules · basic
- ES modules (ESM) · basic
- npm & package management · basic
- Asynchronous JavaScript in Node · basic
- The event loop · basic
- Essential core utilities · basic
- process & configuration · basic
- File system basics · basic
- HTTP & HTTPS servers · medium
- Streams · medium
- Events & EventEmitter · medium
- Advanced filesystem · medium
- crypto · medium
- Compression & encoding · medium
- Child processes · medium
- net, dgram & DNS · medium
- readline, timers & scheduling · medium
- Testing & diagnostics (intro) · medium
- Worker threads · advance
- cluster & multi-process scaling · advance
- Performance & tuning · advance
- Debugging & observability · advance
- Security hardening · advance
- Native addons & N-API · advance
- Architecture patterns · advance
- Graceful shutdown · advance
- 100 Questions · interview questions
- MongoDB — Documents & data modeling (MongoDB)
- Introduction · basic
- Shell, Compass & tools · basic
- Databases & collections · basic
- CRUD operations · basic
- Indexes deep dive · medium
- Explain plans & performance · medium
- Aggregation framework · medium
- Schema design patterns · medium
- Mongoose basics · medium
- Mongoose advanced · medium
- Drivers & connection · medium
- Operators for updates & arrays · medium
- Replication & read preferences · advance
- Write concern & read concern · advance
- Multi-document transactions · advance
- Change streams · advance
- Sharding (overview) · advance
- Atlas Search & full-text · advance
- GridFS & large files · advance
- Backup, restore & ops · advance
- AWS Crash Course (AWS)
- What is Cloud ? · basic
- What is AWS ? · basic
- If not cloud ? · basic
- Cloud Computing · basic
- AWS Pricing · basic
- AWS Shared Responsibility Model · basic
- AWS Management Console · basic
- AWS SDKs · basic
- AWS IAM · medium
- Users, Groups, Roles · medium
- Policies · medium
- AWS Organizations · medium
- AWS Cognito · medium
- AWS Directory Service · medium
- AWS KMS (Key Management Service) · medium
- AWS Secrets Manager · medium
- AWS Shield · medium
- AWS WAF · medium
- AWS Inspector · medium
- AWS GuardDuty · medium
- EC2 · advance
- Launching EC2 Instances · advance
- EBS Volumes · advance
- Security Groups · advance
- Key Pairs · advance
- Elastic IP · advance
- User Data Scripts · advance
- Auto Scaling · advance
- Load Balancers · advance
- ALB · advance
- NLB · advance
- Serverless Compute ,AWS Lambda, Lambda Layers · advance
- Event-Driven Architecture · advance
- ECS · advance
- EKS · advance
#LI-Fulltime
#LI-Hybrid
At Astronomer, we value diversity. We are an equal opportunity employer: we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.