Job Description
The position is for a Senior Cloud Infrastructure & SRE Engineer responsible for ensuring the stability and reliability of core trading systems. This role involves managing AWS and Kubernetes infrastructure, implementing observability systems, and driving automation to enhance operational efficiency.
Key Responsibilities
- Ensure 24/7 stability of core trading systems, including transaction chains, market data services, and related infrastructure. Participate in incident response, troubleshooting, and post-mortems to guarantee service continuity and asset security.
- Design, build, and optimize AWS and EKS/Kubernetes production infrastructure, covering cloud architecture, networking, high availability, capacity planning, and disaster recovery. Manage Kubernetes clusters, including upgrades, node lifecycle, scheduling, resource management, Ingress/LB, DNS, storage, and auto-scaling.
- Develop and enhance observability and stability frameworks (Metrics, Logs, Tracing, Alerting, SLI/SLO, Incident Response, Postmortem) to improve issue detection, diagnosis, and recovery. Collaborate with development teams to address system bottlenecks and architectural risks.
- Promote Infrastructure as Code (IaC), CI/CD, and automation (Terraform/Terragrunt, GitHub Actions, Ansible) to standardize and automate infrastructure and application delivery. Improve change management, release, validation, and rollback processes.
- Manage production operations for core data and middleware services like MySQL, Redis, and Kafka, including HA, monitoring, capacity planning, backup/recovery, performance tuning, and troubleshooting.
- Drive AWS and Kubernetes security governance, covering IAM/RBAC, network isolation, container/image security, secrets management (Secrets/KMS), vulnerability management, and security audits. Coordinate security incident responses with relevant teams.
Job Requirements
- Bachelor’s degree in Computer Science or related field; 5+ years of DevOps/SRE/cloud infrastructure experience. Hands-on experience in large-scale production environments, infrastructure design, and complex incident resolution. Strong risk awareness and cross-team collaboration skills.
- Solid foundation in Linux, networking, and container technologies. Proven expertise in Kubernetes/EKS cluster setup, upgrades, tuning, and troubleshooting (Pod/Node/network level). Familiarity with Java/Go application environments and debugging.
- AWS production experience with core services (IAM, VPC, EC2, EKS, ALB/NLB, Route 53, S3, CloudWatch). Understanding of multi-AZ, HA, network isolation, least privilege, auto-scaling, capacity planning, and cost optimization. Ability to design infrastructure, assess risks, and resolve issues.
- Strong IaC/CI/CD/automation skills (Terraform/Terragrunt, GitHub Actions, Ansible). Proficient in scripting (Bash/Python) to enhance infrastructure and deployment reliability.
- Observability and stability governance expertise (Prometheus, Grafana, ELK/Loki, OpenTelemetry). Knowledge of Metrics/Logs/Tracing, SLI/SLO, alerting, capacity planning, and disaster recovery.
- Production operations experience with MySQL/Redis/Kafka. Security/DevSecOps background (IAM, RBAC, network/container security, Secrets/KMS, vulnerability management, audits). Ability to identify and mitigate security risks.
- Web3/Crypto industry experience preferred, with understanding of blockchain mechanics and digital asset trading risks.
Preferred Qualifications
- Experience in high-availability real-time systems (e.g., crypto exchanges, finance, payments).
- Large-scale/high-traffic Kubernetes/EKS environments, AWS multi-account governance, cross-Region DR, or disaster recovery drills.
- Advanced cloud-native/security/cost optimization practices (Karpenter, Argo CD/Flux, OpenTelemetry, eBPF, FinOps).
Benefits
- Highly competitive salary.
- Singapore work visa sponsorship.
- Flat hierarchy with a diverse, inclusive team culture.
- Generous leave policy (annual/sick), weekends off, and excellent work environment.
Apply directly: https://davionlabs.bamboohr.com/careers/75?source=aWQ9MzM%3D