Job Description
Position Title: Senior Cloud Infrastructure & SRE Engineer
As a Senior Cloud Infrastructure & SRE Engineer, you will play a critical role in ensuring the stability and reliability of our core trading systems. This position requires a proactive approach to managing production environments, optimizing cloud infrastructure, and implementing best practices for observability and automation.
Key Responsibilities
- Ensure 24/7 production stability for core trading services, including trading systems, market data services, and related infrastructure. Participate in major incident response, troubleshooting, risk mitigation, service restoration, and post-mortem analysis to guarantee trading continuity and user asset security.
- Design, implement, and continuously optimize AWS and EKS/Kubernetes production infrastructure, covering cloud architecture, networking, high availability, capacity planning, and disaster recovery. Manage Kubernetes cluster upgrades, node lifecycle, scheduling, resource management, Ingress/Load Balancing, DNS, storage, and auto-scaling.
- Develop and enhance observability and production stability systems (Metrics, Logs, Tracing, Alerting, SLI/SLO, Incident Response, Postmortem) to improve issue detection, diagnosis, and resolution efficiency. Collaborate with development teams to address system bottlenecks, single points of failure, and long-term architectural stability challenges.
- Drive Infrastructure as Code (IaC), CI/CD, and automation initiatives (Terraform/Terragrunt, GitHub Actions, Ansible) to standardize and automate infrastructure and application delivery. Implement robust change management processes including review, deployment, verification, and rollback mechanisms to minimize production risks.
- Manage production operations for core data and middleware services including MySQL, Redis, and Kafka, ensuring high availability, monitoring, capacity planning, backup/recovery, performance optimization, and incident resolution.
- Lead AWS, Kubernetes, and production environment security governance, implementing least privilege access controls (IAM/RBAC), network segmentation, container/image security, secrets management (Secrets/KMS), vulnerability management, and security audits. Coordinate security incident response with relevant teams.
Job Requirements
- Bachelor's degree or higher in Computer Science or related field with 5+ years of DevOps, SRE, or cloud infrastructure experience. Proven track record in large-scale production environment operations, change management, and complex incident resolution. Strong infrastructure design and implementation capabilities with excellent risk awareness and cross-team collaboration skills.
- Solid foundation in Linux, networking, and container technologies with hands-on Kubernetes/EKS experience. Ability to independently manage cluster setup, upgrades, tuning, and troubleshooting at Pod, Node, network, and cluster levels. Familiarity with Java/Go application environments and basic troubleshooting methods.
- Practical AWS production environment experience with core services (IAM, VPC, EC2, EKS, ALB/NLB, Route 53, S3, CloudWatch). Understanding of multi-AZ, high availability, network isolation, least privilege, auto-scaling, capacity planning, and cost optimization principles. Capable of infrastructure design, risk assessment, and incident resolution.
- Strong IaC, CI/CD, and automation skills with experience in Terraform/Terragrunt, GitHub Actions, Ansible. Proficient in scripting (Bash, Python) to enhance infrastructure and application delivery reliability.
- Production observability and stability management experience with Prometheus, Grafana, ELK/Loki, OpenTelemetry. Understanding of Metrics, Logs, Tracing, SLI/SLO, alert management, capacity planning, and disaster recovery mechanisms.
- Operational knowledge of MySQL, Redis, Kafka and security/DevSecOps experience (AWS IAM, Kubernetes RBAC, network/container security, Secrets/KMS, vulnerability management, security audits). Ability to identify and address production security risks.
- Web3/Crypto industry experience preferred, with understanding of blockchain fundamentals and ability to address stability/security risks specific to digital asset trading.
Preferred Qualifications
- Experience with digital asset exchanges, securities, financial trading, or payment systems with high-availability requirements.
- Experience managing large-scale/high-traffic Kubernetes/EKS environments, AWS multi-account governance, cross-Region DR, or disaster recovery drills.
- Advanced cloud-native, cloud security, or cost optimization experience (Karpenter, Argo CD/Flux, OpenTelemetry, eBPF, FinOps).
Benefits
- Highly competitive compensation package
- Singapore work visa sponsorship
- Flat organizational structure with diverse, inclusive team culture
- Generous annual/medical leave policy with weekends off and excellent work environment
Apply directly: https://davionlabs.bamboohr.com/careers/75?source=aWQ9MzM%3D