Staff Software Engineer, Observability
astronomer · Remote
Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications. Trusted by more than 800 of the world's leading enterprises, Astronomer lets businesses do more with their data. To learn more, visit www.astronomer.io .
## About this role:
At Astronomer, our R&D organization is dedicated to providing an exceptional experience in operating data orchestration based on Apache Airflow at many of the world’s largest companies. As we continue to grow, we are looking to solve highly complex technology problems through inventive algorithms and world-class practices around scale.
We are building out a world-class Observability team to deliver data observability capabilities that provide our customers visibility, reliability, and actionable insights into their data pipelines and products across some of the world’s largest enterprises.
We are seeking an experienced Staff Software Engineer to join this effort. In this role, you will contribute directly to the design, development, and scaling of Astronomer’s observability platform, collaborating closely with product, design, and cross-functional engineering teams. Your work will have a direct impact on our customers and the future of the platform.
What you get to do:
• Lead the end-to-end architecture and evolution of major platform components, making foundational design decisions that will shape the long-term direction of the observability platform.
• Build scalable, reliable, and performant features that provide visibility into customer data pipelines.
• Collaborate closely with product, design, and engineering teams to define and deliver high-impact initiatives.
• Write high-quality, maintainable code while following best practices in testing, CI/CD, and operational excellence.
• Set and champion engineering standards for code quality, testing, design, and operational excellence across the team.
• Help improve and evolve the observability platform’s tooling, infrastructure, and processes.
• Lead and coordinate responses to complex production incidents, guiding the team towards rapid resolution and long-term improvements.
• Proactively identify technical opportunities, risks, and gaps - and drive initiatives to address them.
• Mentor engineers of all levels through design reviews, pairing, and feedback, and contribute to the professional growth of the team.
What you bring to the role:
• Strong technical experience building complex distributed systems, preferably with Golang, Python, Kubernetes, SQL/OLAP databases, Kafka, and stream processing.
• Hands-on experience designing, developing, and scaling production infrastructure.
• Excellent collaboration and communication skills to work effectively across teams.
• Focus on delivering customer value through thoughtful technical solutions.
• Enthusiasm for contributing to a healthy, inclusive engineering culture and supporting the growth of teammates.
• Comfort working in areas of high ambiguity, able to break down open-ended problems and drive them to completion with clarity and structure.
Bonus points if you have:
• Experience building or scaling data observability products.
• Familiarity with Apache Airflow or similar workflow orchestrators.
• Background in observability, monitoring, or infrastructure at scale (e.g., Datadog, Honeycomb, or similar).
• Experience working with product teams directly
The estimated salary for this role ranges from
75,000 -
77,000 based on leveling and geography, along with an equity component and a comprehensive benefits package. This range is merely an estimate; actual compensation may deviate from this range based on skills, experience, and qualifications.## About this role:
At Astronomer, our R&D organization is dedicated to providing an exceptional experience in operating data orchestration based on Apache Airflow at many of the world’s largest companies. As we continue to grow, we are looking to solve highly complex technology problems through inventive algorithms and world-class practices around scale.
We are building out a world-class Observability team to deliver data observability capabilities that provide our customers visibility, reliability, and actionable insights into their data pipelines and products across some of the world’s largest enterprises.
We are seeking an experienced Staff Software Engineer to join this effort. In this role, you will contribute directly to the design, development, and scaling of Astronomer’s observability platform, collaborating closely with product, design, and cross-functional engineering teams. Your work will have a direct impact on our customers and the future of the platform.
What you get to do:
• Lead the end-to-end architecture and evolution of major platform components, making foundational design decisions that will shape the long-term direction of the observability platform.
• Build scalable, reliable, and performant features that provide visibility into customer data pipelines.
• Collaborate closely with product, design, and engineering teams to define and deliver high-impact initiatives.
• Write high-quality, maintainable code while following best practices in testing, CI/CD, and operational excellence.
• Set and champion engineering standards for code quality, testing, design, and operational excellence across the team.
• Help improve and evolve the observability platform’s tooling, infrastructure, and processes.
• Lead and coordinate responses to complex production incidents, guiding the team towards rapid resolution and long-term improvements.
• Proactively identify technical opportunities, risks, and gaps - and drive initiatives to address them.
• Mentor engineers of all levels through design reviews, pairing, and feedback, and contribute to the professional growth of the team.
What you bring to the role:
• Strong technical experience building complex distributed systems, preferably with Golang, Python, Kubernetes, SQL/OLAP databases, Kafka, and stream processing.
• Hands-on experience designing, developing, and scaling production infrastructure.
• Excellent collaboration and communication skills to work effectively across teams.
• Focus on delivering customer value through thoughtful technical solutions.
• Enthusiasm for contributing to a healthy, inclusive engineering culture and supporting the growth of teammates.
• Comfort working in areas of high ambiguity, able to break down open-ended problems and drive them to completion with clarity and structure.
Bonus points if you have:
• Experience building or scaling data observability products.
• Familiarity with Apache Airflow or similar workflow orchestrators.
• Background in observability, monitoring, or infrastructure at scale (e.g., Datadog, Honeycomb, or similar).
• Experience working with product teams directly
The estimated salary for this role ranges from
Careeroza — One-stop Zone for Aspirants
Study material, Careeroza mentorship, tech jobs, and career guidance on careeroza.com.
Public study materials
- Django (Django)
- What is Django · basic
- Installing Django · basic
- Features of Django · basic
- MVT Architecture · basic
- Django vs Flask · basic
- Creating Project & Creating App · basic
- Django Project Structure · basic
- URL Routing · basic
- Views · basic
- Templates · basic
- Static & Media Files · medium
- Models · medium
- ORM (Object Relational Mapping) · medium
- Model Relationships · medium
- Migrations · medium
- Django Admin · medium
- Forms · medium
- Authentication · medium
- Authorization · medium
- Middleware · medium
- Signals · medium
- Class Based Views Deep Dive · advance
- Generic Views · advance
- File Handling · advance
- Django REST Framework (DRF) · advance
- Advanced ORM · advance
- Caching · advance
- Asynchronous Django · advance
- Background Tasks · advance
- Interview Questions · interview-questions
- Python (Python)
- Python Fundamentals · basic
- Control Flow · basic
- Strings · basic
- Collections / Data Structures · basic
- Functions · basic
- Modules and Packages · basic
- File Handling · basic
- Exception Handling · basic
- Object-Oriented Programming (OOP · medium
- Advanced Python Concepts · advance
- Functional Programming · advance
- Multithreading & Multiprocessing · advance
- Async Programming · advance
- System Architecture (System Architecture)
- Fundamentals of System Architecture · basic
- Distributed System Basics · basic
- System Reliability Concepts · basic
- Scaling Concepts · basic
- Networking Basics · basic
- Web Communication · basic
- API Communication · basic
- Proxy & Delivery Systems · basic
- Web Architecture Basics · basic
- Rendering Architectures · basic
- Frontend Advanced Concepts · basic
- Message Queue Basics · medium
- What is Load Balancer · medium
- Load Balancing Algorithms · medium
- API Design Basics · medium
- API Protection · medium
- Authentication Basics · medium
- Security Tokens · medium
- Security Threats · medium
- Encryption & Security · medium
- SQL Database Basics · medium
- SQL Scaling Concepts · medium
- NoSQL Databases · medium
- Database Optimization · medium
- Replication Strategies · medium
- Caching Basics · medium
- Cache Storage Systems · medium
- Cache Strategies · medium
- Event-Driven Systems · advance
- Queue Reliability · advance
- Microservices Basics · advance
- Microservice Communication · advance
- Distributed Transactions · advance
- DevOps Basics · advance
- Automation Tools · advance
- Deployment Strategies · advance
- Monitoring Basics · advance
- Monitoring Tools · advance
- Distributed System Concepts · advance
- Distributed Algorithms · advance
- Express.js — Web APIs & middleware (expressjs)
- Application setup · basic
- Routing deep dive · basic
- 1. MVC / Layered Architecture · medium
- Validation · medium
- File uploads · medium
- Sessions & auth (stateful) · medium
- Passport & strategies · medium
- Templating & SSR · medium
- WebSockets & SSE · medium
- Security middleware · advance
- Reverse proxies & trust · advance
- Performance · advance
- API design & versioning · advance
- Testing with Supertest · advance
- GraphQL & tRPC (overview) · advance
- Deployment checklist · advance
- Middlewares · basic
- Request & Response · basic
- System Designing (System Designing)
- Day-1 : What is system Designing ? · basic
- Day-2 : Vertical vs. Horizontal Scaling · basic
- Day-3:How to do vertical scaling ? · basic
- Day4:How to do horizaontal scaling ? · basic
- Day:5TCP vs UDP · basic
- Day6:IP & DNS · basic
- Day7:Client-Server Model · basic
- Day8:HTTP & HTTPS · basic
- Databases (SQL vs NoSQL) · medium
- Caching · medium
- Day9:Latency & Throughput · basic
- Load Balancing · medium
- Indexes & Query Optimization · medium
- CDN · medium
- Proxies · medium
- Message Queues · medium
- Horizontal vs Vertical Scaling · medium
- Database Replication · advance
- Database Sharding · advance
- Consistent Hashing · advance
- CAP Theorem · advance
- Rate Limiting · advance
- Service Discovery · advance
- Event-Driven Architecture · advance
- API Gateway · advance
- Distributed Consensus · expert
- Microservices · expert
- Observability · expert
- Idempotency · expert
- PACELC Theorem · expert
- Two-Phase Commit · expert
- Back-of-Envelope Estimation · expert
- Designing for Failure · expert
- SQL (SQL)
- SQL Fundamentals · basic
- Database Operations · basic
- Table Operations · basic
- CRUD Operations · basic
- Filtering & Operators · basic
- SQL Functions · basic
- GROUPING Data · basic
- Joins · medium
- Constraints · basic
- Subqueries · medium
- Set Operators · medium
- Views · medium
- Indexes · medium
- Normalization · advance
- Transactions · advance
- Stored Procedures & Functions · advance
- Triggers · advance
- Advanced SQL · advance
- Query Optimization · advance
- Database Design · advance
- SQL Security · advance
- Backup & Recovery · advance
- Questions · interview questions
- JavaScript (JavaScript)
- JS Introduction · basic
- Variables & Data Types · basic
- Operators · basic
- Control Flow · basic
- Functions · basic
- Scope & Execution · basic
- Closures · basic
- Objects · basic
- Arrays · basic
- Strings · basic
- DOM Manipulation · basic
- Browser APIs · medium
- Asynchronous JavaScript · medium
- Fetch & APIs · medium
- ES6+ Features · medium
- OOP in JavaScript · medium
- Prototype & Inheritance · advance
- Advanced Functions · advance
- Memory Management · advance
- Error Handling · advance
- Modules · advance
- Advanced Async Concepts · advance
- Functional Programming · advance
- JavaScript Internals · advance
- Performance Optimization · advance
- Angular (Angular)
- Angular Fundamentals · basic
- Project Structure · basic
- Components & Templates · basic
- Data Binding · basic
- Directives · basic
- Pipes · basic
- Component Communication · basic
- Lifecycle Hooks · basic
- Routing Basics · basic
- Routing · basic
- API Calls · basic
- Forms · medium
- Routing · medium
- Services & Dependency Injection · medium
- RxJS & Observables · medium
- Authentication & Security · medium
- Component Interaction · medium
- State Management Basics · medium
- Error Handiling · medium
- Perfomance Basic · medium
- Real World Features · medium
- Advanced Angular Architecture · advance
- Change Detection · advance
- Advanced RxJS · advance
- State Management · advance
- Dynamic Rendering · advance
- Perfomance Optimization · advance
- Modern Angular · advance
- STAR (Situation, Task, Action, and Result) (Situation Based Questions)
- Node.js — Server-side JavaScript (nodejs)
- Getting started · basic
- JavaScript on the server · basic
- CommonJS modules · basic
- ES modules (ESM) · basic
- npm & package management · basic
- Asynchronous JavaScript in Node · basic
- The event loop · basic
- Essential core utilities · basic
- process & configuration · basic
- File system basics · basic
- HTTP & HTTPS servers · medium
- Streams · medium
- Events & EventEmitter · medium
- Advanced filesystem · medium
- crypto · medium
- Compression & encoding · medium
- Child processes · medium
- net, dgram & DNS · medium
- readline, timers & scheduling · medium
- Testing & diagnostics (intro) · medium
- Worker threads · advance
- cluster & multi-process scaling · advance
- Performance & tuning · advance
- Debugging & observability · advance
- Security hardening · advance
- Native addons & N-API · advance
- Architecture patterns · advance
- Graceful shutdown · advance
- 100 Questions · interview questions
- MongoDB — Documents & data modeling (MongoDB)
- Introduction · basic
- Shell, Compass & tools · basic
- Databases & collections · basic
- CRUD operations · basic
- Indexes deep dive · medium
- Explain plans & performance · medium
- Aggregation framework · medium
- Schema design patterns · medium
- Mongoose basics · medium
- Mongoose advanced · medium
- Drivers & connection · medium
- Operators for updates & arrays · medium
- Replication & read preferences · advance
- Write concern & read concern · advance
- Multi-document transactions · advance
- Change streams · advance
- Sharding (overview) · advance
- Atlas Search & full-text · advance
- GridFS & large files · advance
- Backup, restore & ops · advance
- AWS Crash Course (AWS)
- What is Cloud ? · basic
- What is AWS ? · basic
- If not cloud ? · basic
- Cloud Computing · basic
- AWS Pricing · basic
- AWS Shared Responsibility Model · basic
- AWS Management Console · basic
- AWS SDKs · basic
- AWS IAM · medium
- Users, Groups, Roles · medium
- Policies · medium
- AWS Organizations · medium
- AWS Cognito · medium
- AWS Directory Service · medium
- AWS KMS (Key Management Service) · medium
- AWS Secrets Manager · medium
- AWS Shield · medium
- AWS WAF · medium
- AWS Inspector · medium
- AWS GuardDuty · medium
- EC2 · advance
- Launching EC2 Instances · advance
- EBS Volumes · advance
- Security Groups · advance
- Key Pairs · advance
- Elastic IP · advance
- User Data Scripts · advance
- Auto Scaling · advance
- Load Balancers · advance
- ALB · advance
- NLB · advance
- Serverless Compute ,AWS Lambda, Lambda Layers · advance
- Event-Driven Architecture · advance
- ECS · advance
- EKS · advance
#LI-Fulltime
#LI-Hybrid
At Astronomer, we value diversity. We are an equal opportunity employer: we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.