Job Description
Position Title: Senior Cloud Infrastructure & SRE Engineer
We are seeking a highly skilled Senior Cloud Infrastructure & SRE Engineer to join our team. This role is critical in ensuring the stability, security, and efficiency of our core trading infrastructure and services.
Key Responsibilities
- Ensure 24/7 production stability for core trading services, including trading systems, market data services, and related infrastructure. Participate in emergency response for critical incidents, perform root cause analysis, and implement preventive measures to guarantee trading continuity and user asset security.
- Design, implement, and optimize AWS and EKS/Kubernetes production infrastructure, covering cloud architecture, networking, high availability, capacity planning, and disaster recovery. Manage Kubernetes cluster upgrades, node lifecycle, scheduling, resource management, Ingress/Load Balancing, DNS, storage, and auto-scaling.
- Develop and enhance observability and production stability systems (Metrics, Logs, Tracing, Alerting, SLI/SLO, Incident Response, Postmortem) to improve issue detection, diagnosis, and resolution efficiency. Collaborate with development teams to address system bottlenecks and architectural stability challenges.
- Drive Infrastructure as Code (IaC), CI/CD, and automation initiatives (Terraform/Terragrunt, GitHub Actions, Ansible) to standardize and automate infrastructure and application delivery. Implement robust change management processes including review, deployment, verification, and rollback mechanisms.
- Manage production operations for core data and middleware services including MySQL, Redis, Kafka, focusing on high availability, monitoring, capacity planning, backup/recovery, performance tuning, and troubleshooting.
- Implement AWS and Kubernetes security governance including least privilege access (IAM/RBAC), network segmentation, container/image security, secrets management (Secrets/KMS), vulnerability management, and security audits. Coordinate security incident response with relevant teams.
Job Requirements
- Bachelor's degree or higher in Computer Science or related field with 5+ years of DevOps, SRE, or cloud infrastructure experience. Proven track record in large-scale production environment operations, change management, and complex incident resolution.
- Strong foundation in Linux, networking, and container technologies with hands-on Kubernetes/EKS experience including cluster setup, upgrades, tuning, and troubleshooting at pod, node, and network levels.
- Extensive AWS production experience with core services (IAM, VPC, EC2, EKS, ALB/NLB, Route 53, S3, CloudWatch). Deep understanding of multi-AZ, HA, network isolation, least privilege, auto-scaling, capacity planning, and cost optimization principles.
- Proficient in IaC, CI/CD, and automation tools (Terraform/Terragrunt, GitHub Actions, Ansible). Strong scripting skills (Bash, Python) to improve infrastructure and application delivery reliability.
- Experience with production observability and stability tools (Prometheus, Grafana, ELK/Loki, OpenTelemetry) including Metrics, Logs, Tracing, SLI/SLO, alert management, capacity planning, and disaster recovery mechanisms.
- Knowledge of MySQL, Redis, Kafka operations and security practices (AWS IAM, Kubernetes RBAC, network security, container security, Secrets/KMS, vulnerability management, security audits).
- Understanding of Web3/Crypto industry operations and blockchain fundamentals, with ability to identify and mitigate stability and security risks in digital asset trading environments.
Preferred Qualifications
- Experience with digital asset exchanges, securities, financial trading, or payment systems with high-availability requirements.
- Hands-on experience with large-scale/high-traffic Kubernetes/EKS environments, AWS multi-account governance, cross-region DR, or disaster recovery drills.
- Advanced cloud-native, cloud security, or cost optimization experience (Karpenter, Argo CD/Flux, OpenTelemetry, eBPF, FinOps).
Benefits
- Highly competitive compensation package
- Singapore work visa sponsorship
- Flat organizational structure with diverse and inclusive team culture
- Generous annual/medical leave policy with excellent work-life balance
Apply directly at: https://davionlabs.bamboohr.com/careers/75?source=aWQ9MzM%3D