Own design and reliability of ML inference infrastructure serving predictive models and LLMs with high availability and low latency.
Build and optimize low-latency streaming pipelines delivering fresh feature data to production ML models.
Improve distributed training infrastructure for large-scale ML processing.
Develop observability tooling to monitor data quality and detect model degradations.
Mentor junior engineers to raise production-grade software standards.
Requirements
5+ years as a software engineer with ownership of distributed systems in production.
Built/operated low-latency data or ML infrastructure (streaming pipelines, online serving systems, or distributed training) at scale.
Mentored engineers and contributed to rising engineering quality through reviews and leadership.
Familiarity with ML platform components (feature stores, model serving frameworks, training orchestration) to partner with ML engineers as a platform builder.
Uses generative AI responsibly with human oversight to deliver business-ready outputs and improve workflow efficiency, cost, and quality.
Compensation & benefits
Base salary $191,250—$225,000 USD (excluding equity and bonus).
Total compensation may include equity and bonus eligibility; benefits include medical, dental, vision, and 401(k).
Remote-first, but not remote-only; quarterly in-person in-work sessions (surges).