Job Description
Position Title: Senior Cloud Infrastructure & SRE Engineer
Key Responsibilities:
- Ensure 24/7 production stability for core trading systems, covering trading workflows, market data services, and related infrastructure. Participate in major incident response, including troubleshooting, risk mitigation, service recovery, and post-mortem analysis to guarantee trading continuity and user asset security.
- Design and optimize AWS/EKS/Kubernetes production infrastructure, including cloud architecture, networking, HA, capacity planning, disaster recovery, cluster upgrades, node lifecycle management, scheduling/resource allocation, Ingress/LB, DNS, storage, and auto-scaling. Continuously improve cloud cost efficiency while contributing to architecture design from stability, security, and operability perspectives.
- Build and enhance observability and stability frameworks (Metrics/Logs/Tracing/Alerting/SLI-SLO/Incident Response/Postmortem) to accelerate issue detection, diagnosis, and resolution. Collaborate with R&D to address systemic bottlenecks, SPOFs, and architectural stability challenges.
- Drive IaC/CI-CD/automation initiatives (Terraform/Terragrunt/GitHub Actions/Ansible) to standardize infrastructure/app delivery processes. Improve change management with rigorous review, deployment, verification, and rollback mechanisms.
- Manage production operations for core data services (MySQL/Redis/Kafka) including HA, monitoring, capacity planning, backup/recovery, performance tuning, and incident handling.
- Implement AWS/K8s security governance covering least-privilege access (IAM/RBAC), network segmentation, container/image security, secrets management (KMS), vulnerability remediation, and security audits. Coordinate security incident response across teams.
Job Requirements
- Bachelor's degree in Computer Science or related field with 5+ years in DevOps/SRE/cloud infrastructure. Proven experience in large-scale production operations, change management, and complex troubleshooting. Strong architectural design and implementation skills with risk awareness and cross-team collaboration abilities.
- Solid foundation in Linux, networking, and container technologies. Hands-on Kubernetes/EKS expertise including cluster deployment/upgrades/tuning and multi-layer troubleshooting (Pod/Node/Network). Familiarity with Java/Go application environments.
- AWS production experience with core services (IAM/VPC/EC2/EKS/ALB/Route53/S3/CloudWatch). Understanding of multi-AZ, HA, network isolation, least privilege, auto-scaling, capacity planning, and cost optimization principles.
- Strong IaC/CI-CD/automation skills using Terraform/GitHub Actions/Ansible. Proficient in scripting (Bash/Python) to enhance infrastructure/app delivery reliability.
- Observability/stability expertise with Prometheus/Grafana/ELK/OpenTelemetry. Knowledge of Metrics/Logs/Tracing/SLI-SLO/alerting/capacity planning/DR mechanisms.
- Operational experience with MySQL/Redis/Kafka and security practices (IAM/RBAC/network/container/host security/KMS/vulnerability management/auditing).
- Web3/Crypto production experience preferred, with understanding of blockchain fundamentals and digital asset trading risks.
Preferred Qualifications:
- Experience in digital asset exchanges, securities, financial trading, or high-availability payment systems.
- Expertise in large-scale/high-traffic Kubernetes/EKS environments, AWS multi-account governance, or cross-Region DR.
- Advanced cloud-native/security/cost optimization skills (Karpenter/ArgoCD/OpenTelemetry/eBPF/FinOps).
Benefits
- Highly competitive compensation package
- Singapore work visa sponsorship
- Flat organizational structure with diverse, inclusive culture
- Generous PTO/sick leave with excellent work-life balance
Apply directly: https://davionlabs.bamboohr.com/careers/75?source=aWQ9MzM%3D