Architect and operate scalable, self-healing infrastructure across multi-region deployments using Kubernetes, Terraform, and cloud-native tools.
Drive AI enablement across engineering with agentic development tools (Claude Code, Cursor, Codex).
Build AI-powered infrastructure tooling and automation (e.g., automated K8s upgrades, IaC plan analysis, cost optimization advisors, MCP servers, n8n workflows).
Build and maintain internal developer platform (IDP) capabilities for self-service deployments, observability, and reliability.
Develop observability frameworks with Prometheus and Grafana for metrics, dashboards, and alerting, and lead incident management with blameless post-mortems; define SLIs/SLOs and error budgets across services.
Requirements
5+ years as an Infrastructure Engineer focused on reliability (SRE/Production/Platform).
Experience driving company-wide reliability efforts, including SLO frameworks and error budget policies.
Strong proficiency with observability stacks: OpenTelemetry, Prometheus/Grafana.
Deep experience with cloud infrastructure (AWS/GCP), Kubernetes, and multi-region architectures.
Skilled with Terraform, Helm, and GitOps workflows (e.g., ArgoCD) with an automation-first mindset.
Compensation & benefits
Competitive base salary plus equity.
Medical, dental, and vision coverage.
401k and unlimited flexible time off.
Gym reimbursement.
Home office build-out budget; in-office group meals; wellbeing & mental health perks; learning and development stipend; company-sponsored conferences & events; fertility benefits; HSA and FSA plans.