Staff Machine Learning Engineer
handshake · Remote
Experience: 8+ years of experience
## About Handshake
Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.
In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~
Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.
In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~
B run rate and pay ~$60M to over 30K individuals every month.
Why join Handshake now:
• Shape how every career evolves in the AI economy, at global scale, with impact your friends, family, and peers can see and feel
• Partner hand-in-hand with world-class AI labs, Fortune 500 partners, and the world's top educational institutions
• Work alongside engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders
• Build a massive, fast-growing business with billions in revenue
## About Handshake AI
Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3–5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.
## About the Role
Handshake is hiring a Staff Machine Learning Engineer for the Network and Handshake AI Marketplace Relevance team. AI is transforming how students navigate their careers, and we're committed to providing innovative, responsible AI-powered solutions that guide students from educational aspirations to meaningful career opportunities.
In this role, you will set the technical direction for the systems that power core embedding models, consumer job search & recommendations, user understanding, and personalized notifications across the Handshake platform. You'll operate as a force multiplier — architecting the ML infrastructure that the broader Relevance org builds on, and raising the technical bar for how the team ships models into production.
You'll own the roadmap for a suite of systems built on multiple retrieval models — Graph Neural Network, bi-encoders, semantic cross-encoders with multi-stage rankers, running on a data platform with billions of data points. In addition, the team is investing into areas of generative retrieval and post-training. Your work will directly move the marketplace's core metrics, and you'll be a key voice in how Handshake approaches explainability, fairness, and quality as we scale responsible AI across the product.
## What You'll Do
Architect: Define the technical strategy and system design for ML models and infrastructure spanning search and recommendation, notifications, generative retrieval, and core embeddings — making build-vs-buy, architecture, and platform decisions with company-wide impact.
Force Multiplier: Set technical standards and best practices for model development, experimentation, and production deployment; mentor and elevate engineers and data scientists across the team.
Cross-functional Leader: Partner with engineering leadership, product, and data science to translate ambiguous business problems into a clear technical roadmap, and drive alignment across stakeholders on priorities and tradeoffs.
Operator: Get hands-on where it matters most — building and shipping the highest-leverage models and systems yourself, and unblocking the team on the hardest technical problems.
## Desired Capabilities
• 8+ years of experience in machine learning, data science, or a related field, with a track record of owning large-scale, production ML systems end-to-end
• Deep expertise in Python and ML frameworks such as scikit-learn, PyTorch, or TensorFlow
• Experience in recommendations, personalization, NLP, deep learning, LLMs, or explainable AI
• Deep familiarity with the ML lifecycle (experiment tracking, model monitoring, feature pipelines) at scale
• Demonstrated ability to architect and scale ML infrastructure — embedding-based retrieval, ranking systems, GNNs, or similar — in a high-traffic cloud based production environment
• Strong foundation in core ML concepts (classification, regression, ranking, model evaluation) with the judgment to know when and how to apply them
• Experience setting technical direction across teams and mentoring senior and mid-level engineers
• Track record of driving measurable business impact through ML systems at scale
• Experience with Generative Retrieval and LLM Post Training recipes is a plus.
## Extra Credit
• A track record as a clear, persuasive communicator who can align technical and non-technical stakeholders around a shared roadmap
• Experience building or scaling a team's technical practices and standards from the ground up
## Perks
Handshake delivers benefits that help you feel supported — and thrive at work and in life.
The below benefits are for full-time US employees.
• 🎯 Ownership: Equity in a fast-growing company
• 💰 Financial Wellness: 401(k) match, competitive compensation, financial coaching
• 🍼 Family Support: Paid parental leave, fertility benefits, parental coaching
• 💝 Wellbeing: Medical, dental, and vision, mental health support, $500 wellness stipend
• 📚 Growth:
,000 learning stipend, ongoing development
• 💻 Remote & Office: Internet, commuting, and free lunch/gym in our SF office
• 🏝 Time Off: Flexible PTO, 15 holidays + 2 flex days
• 🤝 Connection: Team outings & referral bonuses
Explore our mission, values, and comprehensive US benefits at joinhandshake.com/careers .
Why join Handshake now:
• Shape how every career evolves in the AI economy, at global scale, with impact your friends, family, and peers can see and feel
• Partner hand-in-hand with world-class AI labs, Fortune 500 partners, and the world's top educational institutions
• Work alongside engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders
• Build a massive, fast-growing business with billions in revenue
## About Handshake AI
Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3–5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.
## About the Role
Handshake is hiring a Staff Machine Learning Engineer for the Network and Handshake AI Marketplace Relevance team. AI is transforming how students navigate their careers, and we're committed to providing innovative, responsible AI-powered solutions that guide students from educational aspirations to meaningful career opportunities.
In this role, you will set the technical direction for the systems that power core embedding models, consumer job search & recommendations, user understanding, and personalized notifications across the Handshake platform. You'll operate as a force multiplier — architecting the ML infrastructure that the broader Relevance org builds on, and raising the technical bar for how the team ships models into production.
You'll own the roadmap for a suite of systems built on multiple retrieval models — Graph Neural Network, bi-encoders, semantic cross-encoders with multi-stage rankers, running on a data platform with billions of data points. In addition, the team is investing into areas of generative retrieval and post-training. Your work will directly move the marketplace's core metrics, and you'll be a key voice in how Handshake approaches explainability, fairness, and quality as we scale responsible AI across the product.
## What You'll Do
Architect: Define the technical strategy and system design for ML models and infrastructure spanning search and recommendation, notifications, generative retrieval, and core embeddings — making build-vs-buy, architecture, and platform decisions with company-wide impact.
Force Multiplier: Set technical standards and best practices for model development, experimentation, and production deployment; mentor and elevate engineers and data scientists across the team.
Cross-functional Leader: Partner with engineering leadership, product, and data science to translate ambiguous business problems into a clear technical roadmap, and drive alignment across stakeholders on priorities and tradeoffs.
Operator: Get hands-on where it matters most — building and shipping the highest-leverage models and systems yourself, and unblocking the team on the hardest technical problems.
## Desired Capabilities
• 8+ years of experience in machine learning, data science, or a related field, with a track record of owning large-scale, production ML systems end-to-end
• Deep expertise in Python and ML frameworks such as scikit-learn, PyTorch, or TensorFlow
• Experience in recommendations, personalization, NLP, deep learning, LLMs, or explainable AI
• Deep familiarity with the ML lifecycle (experiment tracking, model monitoring, feature pipelines) at scale
• Demonstrated ability to architect and scale ML infrastructure — embedding-based retrieval, ranking systems, GNNs, or similar — in a high-traffic cloud based production environment
• Strong foundation in core ML concepts (classification, regression, ranking, model evaluation) with the judgment to know when and how to apply them
• Experience setting technical direction across teams and mentoring senior and mid-level engineers
• Track record of driving measurable business impact through ML systems at scale
• Experience with Generative Retrieval and LLM Post Training recipes is a plus.
## Extra Credit
• A track record as a clear, persuasive communicator who can align technical and non-technical stakeholders around a shared roadmap
• Experience building or scaling a team's technical practices and standards from the ground up
## Perks
Handshake delivers benefits that help you feel supported — and thrive at work and in life.
The below benefits are for full-time US employees.
• 🎯 Ownership: Equity in a fast-growing company
• 💰 Financial Wellness: 401(k) match, competitive compensation, financial coaching
• 🍼 Family Support: Paid parental leave, fertility benefits, parental coaching
• 💝 Wellbeing: Medical, dental, and vision, mental health support, $500 wellness stipend
• 📚 Growth:
Careeroza — One-stop Zone for Aspirants
Study material, Careeroza mentorship, tech jobs, and career guidance on careeroza.com.
Public study materials
- Django (Django)
- What is Django · basic
- Installing Django · basic
- Features of Django · basic
- MVT Architecture · basic
- Django vs Flask · basic
- Creating Project & Creating App · basic
- Django Project Structure · basic
- URL Routing · basic
- Views · basic
- Templates · basic
- Static & Media Files · medium
- Models · medium
- ORM (Object Relational Mapping) · medium
- Model Relationships · medium
- Migrations · medium
- Django Admin · medium
- Forms · medium
- Authentication · medium
- Authorization · medium
- Middleware · medium
- Signals · medium
- Class Based Views Deep Dive · advance
- Generic Views · advance
- File Handling · advance
- Django REST Framework (DRF) · advance
- Advanced ORM · advance
- Caching · advance
- Asynchronous Django · advance
- Background Tasks · advance
- Interview Questions · interview-questions
- Python (Python)
- Python Fundamentals · basic
- Control Flow · basic
- Strings · basic
- Collections / Data Structures · basic
- Functions · basic
- Modules and Packages · basic
- File Handling · basic
- Exception Handling · basic
- Object-Oriented Programming (OOP · medium
- Advanced Python Concepts · advance
- Functional Programming · advance
- Multithreading & Multiprocessing · advance
- Async Programming · advance
- System Architecture (System Architecture)
- Fundamentals of System Architecture · basic
- Distributed System Basics · basic
- System Reliability Concepts · basic
- Scaling Concepts · basic
- Networking Basics · basic
- Web Communication · basic
- API Communication · basic
- Proxy & Delivery Systems · basic
- Web Architecture Basics · basic
- Rendering Architectures · basic
- Frontend Advanced Concepts · basic
- Message Queue Basics · medium
- What is Load Balancer · medium
- Load Balancing Algorithms · medium
- API Design Basics · medium
- API Protection · medium
- Authentication Basics · medium
- Security Tokens · medium
- Security Threats · medium
- Encryption & Security · medium
- SQL Database Basics · medium
- SQL Scaling Concepts · medium
- NoSQL Databases · medium
- Database Optimization · medium
- Replication Strategies · medium
- Caching Basics · medium
- Cache Storage Systems · medium
- Cache Strategies · medium
- Event-Driven Systems · advance
- Queue Reliability · advance
- Microservices Basics · advance
- Microservice Communication · advance
- Distributed Transactions · advance
- DevOps Basics · advance
- Automation Tools · advance
- Deployment Strategies · advance
- Monitoring Basics · advance
- Monitoring Tools · advance
- Distributed System Concepts · advance
- Distributed Algorithms · advance
- Express.js — Web APIs & middleware (expressjs)
- Application setup · basic
- Routing deep dive · basic
- 1. MVC / Layered Architecture · medium
- Validation · medium
- File uploads · medium
- Sessions & auth (stateful) · medium
- Passport & strategies · medium
- Templating & SSR · medium
- WebSockets & SSE · medium
- Security middleware · advance
- Reverse proxies & trust · advance
- Performance · advance
- API design & versioning · advance
- Testing with Supertest · advance
- GraphQL & tRPC (overview) · advance
- Deployment checklist · advance
- Middlewares · basic
- Request & Response · basic
- System Designing (System Designing)
- Day-1 : What is system Designing ? · basic
- Day-2 : Vertical vs. Horizontal Scaling · basic
- Day-3:How to do vertical scaling ? · basic
- Day4:How to do horizaontal scaling ? · basic
- Day:5TCP vs UDP · basic
- Day6:IP & DNS · basic
- Day7:Client-Server Model · basic
- Day8:HTTP & HTTPS · basic
- Databases (SQL vs NoSQL) · medium
- Caching · medium
- Day9:Latency & Throughput · basic
- Load Balancing · medium
- Indexes & Query Optimization · medium
- CDN · medium
- Proxies · medium
- Message Queues · medium
- Horizontal vs Vertical Scaling · medium
- Database Replication · advance
- Database Sharding · advance
- Consistent Hashing · advance
- CAP Theorem · advance
- Rate Limiting · advance
- Service Discovery · advance
- Event-Driven Architecture · advance
- API Gateway · advance
- Distributed Consensus · expert
- Microservices · expert
- Observability · expert
- Idempotency · expert
- PACELC Theorem · expert
- Two-Phase Commit · expert
- Back-of-Envelope Estimation · expert
- Designing for Failure · expert
- SQL (SQL)
- SQL Fundamentals · basic
- Database Operations · basic
- Table Operations · basic
- CRUD Operations · basic
- Filtering & Operators · basic
- SQL Functions · basic
- GROUPING Data · basic
- Joins · medium
- Constraints · basic
- Subqueries · medium
- Set Operators · medium
- Views · medium
- Indexes · medium
- Normalization · advance
- Transactions · advance
- Stored Procedures & Functions · advance
- Triggers · advance
- Advanced SQL · advance
- Query Optimization · advance
- Database Design · advance
- SQL Security · advance
- Backup & Recovery · advance
- Questions · interview questions
- JavaScript (JavaScript)
- JS Introduction · basic
- Variables & Data Types · basic
- Operators · basic
- Control Flow · basic
- Functions · basic
- Scope & Execution · basic
- Closures · basic
- Objects · basic
- Arrays · basic
- Strings · basic
- DOM Manipulation · basic
- Browser APIs · medium
- Asynchronous JavaScript · medium
- Fetch & APIs · medium
- ES6+ Features · medium
- OOP in JavaScript · medium
- Prototype & Inheritance · advance
- Advanced Functions · advance
- Memory Management · advance
- Error Handling · advance
- Modules · advance
- Advanced Async Concepts · advance
- Functional Programming · advance
- JavaScript Internals · advance
- Performance Optimization · advance
- Angular (Angular)
- Angular Fundamentals · basic
- Project Structure · basic
- Components & Templates · basic
- Data Binding · basic
- Directives · basic
- Pipes · basic
- Component Communication · basic
- Lifecycle Hooks · basic
- Routing Basics · basic
- Routing · basic
- API Calls · basic
- Forms · medium
- Routing · medium
- Services & Dependency Injection · medium
- RxJS & Observables · medium
- Authentication & Security · medium
- Component Interaction · medium
- State Management Basics · medium
- Error Handiling · medium
- Perfomance Basic · medium
- Real World Features · medium
- Advanced Angular Architecture · advance
- Change Detection · advance
- Advanced RxJS · advance
- State Management · advance
- Dynamic Rendering · advance
- Perfomance Optimization · advance
- Modern Angular · advance
- STAR (Situation, Task, Action, and Result) (Situation Based Questions)
- Node.js — Server-side JavaScript (nodejs)
- Getting started · basic
- JavaScript on the server · basic
- CommonJS modules · basic
- ES modules (ESM) · basic
- npm & package management · basic
- Asynchronous JavaScript in Node · basic
- The event loop · basic
- Essential core utilities · basic
- process & configuration · basic
- File system basics · basic
- HTTP & HTTPS servers · medium
- Streams · medium
- Events & EventEmitter · medium
- Advanced filesystem · medium
- crypto · medium
- Compression & encoding · medium
- Child processes · medium
- net, dgram & DNS · medium
- readline, timers & scheduling · medium
- Testing & diagnostics (intro) · medium
- Worker threads · advance
- cluster & multi-process scaling · advance
- Performance & tuning · advance
- Debugging & observability · advance
- Security hardening · advance
- Native addons & N-API · advance
- Architecture patterns · advance
- Graceful shutdown · advance
- 100 Questions · interview questions
- MongoDB — Documents & data modeling (MongoDB)
- Introduction · basic
- Shell, Compass & tools · basic
- Databases & collections · basic
- CRUD operations · basic
- Indexes deep dive · medium
- Explain plans & performance · medium
- Aggregation framework · medium
- Schema design patterns · medium
- Mongoose basics · medium
- Mongoose advanced · medium
- Drivers & connection · medium
- Operators for updates & arrays · medium
- Replication & read preferences · advance
- Write concern & read concern · advance
- Multi-document transactions · advance
- Change streams · advance
- Sharding (overview) · advance
- Atlas Search & full-text · advance
- GridFS & large files · advance
- Backup, restore & ops · advance
- AWS Crash Course (AWS)
- What is Cloud ? · basic
- What is AWS ? · basic
- If not cloud ? · basic
- Cloud Computing · basic
- AWS Pricing · basic
- AWS Shared Responsibility Model · basic
- AWS Management Console · basic
- AWS SDKs · basic
- AWS IAM · medium
- Users, Groups, Roles · medium
- Policies · medium
- AWS Organizations · medium
- AWS Cognito · medium
- AWS Directory Service · medium
- AWS KMS (Key Management Service) · medium
- AWS Secrets Manager · medium
- AWS Shield · medium
- AWS WAF · medium
- AWS Inspector · medium
- AWS GuardDuty · medium
- EC2 · advance
- Launching EC2 Instances · advance
- EBS Volumes · advance
- Security Groups · advance
- Key Pairs · advance
- Elastic IP · advance
- User Data Scripts · advance
- Auto Scaling · advance
- Load Balancers · advance
- ALB · advance
- NLB · advance
- Serverless Compute ,AWS Lambda, Lambda Layers · advance
- Event-Driven Architecture · advance
- ECS · advance
- EKS · advance
• 💻 Remote & Office: Internet, commuting, and free lunch/gym in our SF office
• 🏝 Time Off: Flexible PTO, 15 holidays + 2 flex days
• 🤝 Connection: Team outings & referral bonuses
Explore our mission, values, and comprehensive US benefits at joinhandshake.com/careers .