Job Description
Position Title: Senior Cloud Infrastructure & SRE Engineer
We are seeking an experienced Senior Cloud Infrastructure & SRE Engineer to join our team. This role is critical in ensuring the stability, security, and efficiency of our core trading platform and related infrastructure.
Key Responsibilities
- Ensure 24/7 production stability for core trading services, including trading systems, market data services, and related infrastructure. Participate in emergency response for critical incidents, including troubleshooting, risk mitigation, service restoration, and post-mortem analysis.
- Design, implement, and optimize AWS and EKS/Kubernetes production infrastructure, covering cloud architecture, networking, high availability, capacity planning, and disaster recovery. Manage Kubernetes cluster upgrades, node lifecycle, scheduling, resource management, Ingress/LB, DNS, storage, and auto-scaling.
- Develop and enhance observability and production stability systems (Metrics, Logs, Tracing, Alerting, SLI/SLO, Incident Response, Postmortem) to improve issue detection, diagnosis, and resolution efficiency.
- Promote Infrastructure as Code (IaC), CI/CD, and automation (Terraform/Terragrunt, GitHub Actions, Ansible) to standardize infrastructure and application delivery processes, while improving change management procedures.
- Manage production operations for core data and middleware services including MySQL, Redis, and Kafka, covering HA, monitoring, capacity planning, backup/recovery, performance tuning, and troubleshooting.
- Implement security governance for AWS, Kubernetes, and production environments, including IAM/RBAC, network isolation, container/image security, secrets management, vulnerability management, and security audits.
Job Requirements
- Bachelor's degree in Computer Science or related field with 5+ years of DevOps/SRE/cloud infrastructure experience. Proven experience in large-scale production operations, change management, and complex incident resolution.
- Strong Linux, networking, and container expertise with hands-on Kubernetes/EKS experience in cluster setup, upgrades, tuning, and troubleshooting. Basic knowledge of Java/Go application environments.
- Extensive AWS production experience with core services (IAM, VPC, EC2, EKS, ALB/NLB, Route53, S3, CloudWatch). Understanding of multi-AZ, HA, network isolation, least privilege, auto-scaling, and capacity planning principles.
- Proficient in IaC, CI/CD, and automation tools (Terraform/Terragrunt, GitHub Actions, Ansible). Strong scripting skills (Bash/Python) for improving infrastructure and application delivery reliability.
- Experience with production observability stacks (Prometheus, Grafana, ELK/Loki, OpenTelemetry) and stability governance (Metrics, Logs, Tracing, SLI/SLO, alerting, capacity planning, DR).
- Knowledge of MySQL/Redis/Kafka operations and security practices (IAM, RBAC, network security, container security, secrets management, vulnerability management).
- Web3/Crypto industry experience preferred, with understanding of blockchain fundamentals and digital asset trading security considerations.
Preferred Qualifications
- Experience in digital asset exchanges, securities, financial trading, or payment systems with high-availability requirements.
- Hands-on experience with large-scale/high-traffic Kubernetes/EKS environments, AWS multi-account governance, or cross-region DR.
- Deep expertise in cloud-native technologies, cloud security, or cost optimization (Karpenter, Argo CD/Flux, OpenTelemetry, eBPF, FinOps).
Benefits
- Highly competitive compensation package
- Singapore work visa sponsorship
- Flat organizational structure with diverse, inclusive team culture
- Generous vacation/sick leave policy with excellent work-life balance
Apply directly: https://davionlabs.bamboohr.com/careers/75?source=aWQ9MzM%3D