Data Operations Engineer, X-Layer
OKX
2 months
Singapore
•
Engineer
Python
SQL
Cloud
Crypto
Kubernetes
Terraform
Spark
AWS
Big Data
Linux
SRE
Grafana
Prometheus
Ansible
Databricks
Platform Engineering
Infrastructure as Code
Singapore
Datadog
GitLab CI/CD
Trino
Data Operations
Alibaba Cloud DataWorks
MaxCompute
ODPS
Hologres
VVP
Flink
StarRocks
Hive
Presto
Shell
CloudWatch
ELK
OpenSearch
Jenkins
OLAP
Cloud Cost Optimization
Resource Governance
Capacity Planning
Data Platform Operations
Big Data Engineering
At OKX, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom. OKX is a leading crypto exchange and the developer of OKX Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). OKX is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. Across our multiple offices globally, we are united by our core principles: We Before Me, Do the Right Thing, and Get Things Done. These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er. OKX is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more. We are hiring Data Operations Engineers / Data Platform SREs in Singapore to support the reliability, observability, performance, and cost efficiency of our multi-cloud data platform. You will work on core data platform components across Alibaba Cloud and AWS, covering daily operations, incident response, monitoring, automation, resource optimization, and platform reliability improvements. You will collaborate closely with data engineering, data warehouse, BI, platform teams, and cloud vendors to ensure our data pipelines and platform services run reliably at scale. What You'll Be Doing: Operate and support core data platform components, including Alibaba Cloud DataWorks, MaxCompute / ODPS, Hologres, VVP / Flink, and AWS-based platforms such as Databricks and StarRocks. Build and maintain monitoring, alerting, and SLA / SLO metrics for data platform services. Respond to production incidents, troubleshoot issues, participate in post-incident reviews, and drive long-term fixes. Analyze compute, storage, job, and cluster resource usage to improve performance and optimize cloud costs. Develop or integrate automation scripts and operational tools to improve health checks, releases, scaling, and troubleshooting efficiency. Collaborate with data engineering, data warehouse, BI, platform teams, and cloud vendors to continuously improve platform reliability and operational efficiency.