Job Description
Position Title: Senior Cloud Infrastructure & SRE Engineer
Key Responsibilities:
- Ensure 24/7 production stability for core trading systems, covering trading links, market data services, and related infrastructure. Participate in major incident response, including troubleshooting, risk mitigation, service recovery, and post-mortem analysis to guarantee trading continuity and user asset security.
- Design and optimize AWS/EKS/Kubernetes production infrastructure, including cloud architecture, networking, high availability, capacity planning, disaster recovery, cluster upgrades, node lifecycle management, scheduling/resource allocation, Ingress/LB, DNS, storage, and auto-scaling. Continuously improve cloud resource efficiency and cost-effectiveness while contributing to architectural design from stability, security, and operability perspectives.
- Build and enhance observability and stability systems (Metrics, Logs, Tracing, Alerting, SLI/SLO, Incident Response, Postmortem) to accelerate problem detection, diagnosis, and resolution. Collaborate with R&D to address systemic bottlenecks, single points of failure, and long-term architectural stability issues.
- Drive IaC, CI/CD, and automation initiatives (Terraform/Terragrunt, GitHub Actions, Ansible) to standardize infrastructure and application delivery. Improve change management processes including review, deployment, verification, and rollback to minimize production risks.
- Manage production operations for core data/middleware services (MySQL, Redis, Kafka), covering HA, monitoring, capacity planning, backup/recovery, performance tuning, and incident handling.
- Implement AWS/Kubernetes security governance including least privilege access (IAM/RBAC), network segmentation, container/image security, secrets management (Secrets/KMS), vulnerability management, and security audits. Coordinate security incident response with relevant teams.
Job Requirements
- Bachelor's degree in Computer Science or related field with 5+ years in DevOps/SRE/cloud infrastructure roles. Proven experience in large-scale production operations, change management, and complex troubleshooting. Strong infrastructure design/implementation skills with risk awareness and cross-team collaboration abilities.
- Solid foundation in Linux, networking, and container technologies. Hands-on Kubernetes/EKS experience including cluster deployment, upgrades, tuning, and troubleshooting at pod/node/network levels. Familiarity with Java/Go application environments and basic debugging.
- AWS production experience with core services (IAM, VPC, EC2, EKS, ALB/NLB, Route53, S3, CloudWatch). Understanding of multi-AZ, HA, network isolation, least privilege, auto-scaling, capacity planning, and cost optimization principles.
- Strong IaC/CI/CD/automation skills using Terraform/Terragrunt, GitHub Actions, Ansible. Proficient in scripting (Bash/Python) to improve infrastructure/app delivery reliability.
- Observability/stability expertise with Prometheus, Grafana, ELK/Loki, OpenTelemetry. Deep understanding of Metrics/Logs/Tracing, SLI/SLO, alert management, capacity planning, and DR mechanisms.
- Production operations experience with MySQL/Redis/Kafka. Security operations/DevSecOps knowledge including IAM/RBAC, network/host/container security, secrets management, vulnerability handling, and security audits.
- Web3/Crypto industry experience preferred, with understanding of blockchain fundamentals and ability to address stability/security risks in digital asset trading systems.
Preferred Qualifications:
- Experience in digital asset exchanges, securities, financial trading, or payment systems with high-availability requirements.
- Large-scale/high-traffic Kubernetes/EKS environments, AWS multi-account governance, cross-Region DR, or disaster recovery drills.
- Advanced cloud-native/security/cost optimization expertise (Karpenter, ArgoCD/Flux, OpenTelemetry, eBPF, FinOps).
Benefits
- Highly competitive compensation package
- Singapore work visa sponsorship
- Flat organizational structure with diverse, inclusive culture
- Generous annual/medical leave, weekends off, and excellent work environment
Apply directly: https://davionlabs.bamboohr.com/careers/75?source=aWQ9MzM%3D