Category Labs (formerly known as Monad Labs) is a team of systems engineers and researchers on a mission to design and build at the frontier of decentralized technology. We strive to deliver significant improvements over existing blockchain solutions. After raising $225M in series A funding, led by Paradigm, we are growing our team. We’re the team behind Monad, a high-performance, EVM-compatible Layer 1 whose public mainnet is now live. We write the core software that runs it: a parallel-execution EVM, a custom state database, and a BFT consensus client, all developed in the open. We're looking for a Senior DevOps / Infrastructure Engineer to operate the infrastructure behind Monad, and to push how much of that operation can be driven by AI. You'll keep our globally-distributed validator, full node, and archive fleet healthy across mainnet and testnet, own our infrastructure-as-code and observability, and build the agentic tooling and guardrails that let a small team safely operate a large fleet. As more of our engineering shifts toward AI, this role is central to designing the workflows, deterministic guardrails, and security perimeters within which autonomous agents operate our infrastructure; you'll also help stand up and operate the infrastructure behind our own growing model workloads. Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response. Own our infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services. Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives. Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes). Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight. Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse. Harden nodes and services, manage secrets, and continuously drive down manual toil.