Job Description
Position Title: Senior Cloud Infrastructure & SRE Engineer
As a Senior Cloud Infrastructure & SRE Engineer, you will play a critical role in ensuring the stability, security, and efficiency of our core trading systems and related infrastructure. This position requires a combination of technical expertise, problem-solving skills, and the ability to work collaboratively across teams to maintain high availability and performance of our production environment.
Key Responsibilities
- Ensure 24/7 stability of core trading systems, including transaction chains, market data services, and related infrastructure. Participate in emergency response for critical incidents, including problem identification, risk control, service recovery, and post-mortem analysis to guarantee trading continuity and user asset security.
- Design, implement, and continuously optimize AWS and EKS/Kubernetes production infrastructure, covering cloud architecture, networking, high availability, capacity planning, and disaster recovery. Manage Kubernetes cluster upgrades, node lifecycle, scheduling, resource management, Ingress/Load Balancing, DNS, storage, and auto-scaling.
- Develop and enhance observability and production stability systems (Metrics, Logs, Tracing, Alerting, SLI/SLO, Incident Response, Postmortem) to improve problem detection, diagnosis, and recovery efficiency. Collaborate with development teams to address system bottlenecks, single points of failure, and long-term architectural stability issues.
- Promote Infrastructure as Code (IaC), CI/CD, and automation practices (Terraform/Terragrunt, GitHub Actions, Ansible) to standardize and automate infrastructure and application delivery. Improve change review, deployment, verification, and rollback mechanisms to minimize production risks.
- Manage production operations for core data and middleware services including MySQL, Redis, and Kafka, ensuring high availability, monitoring, capacity planning, backup/recovery, performance optimization, and incident resolution.
- Implement security governance for AWS, Kubernetes, and production environments, including least privilege access control (IAM/RBAC), network isolation, container/image security, secrets management (Secrets/KMS), vulnerability management, and security audits. Coordinate with relevant teams for security incident response.
Job Requirements
- Bachelor's degree or higher in Computer Science or related field with 5+ years of DevOps, SRE, or cloud infrastructure experience. Proven experience in large-scale production environment operations, change management, and complex incident resolution. Strong infrastructure design and implementation capabilities with excellent risk awareness and cross-team collaboration skills.
- Solid foundation in Linux, networking, and container technologies with hands-on production Kubernetes/EKS experience. Ability to independently perform cluster setup, upgrades, tuning, and troubleshooting at Pod, Node, network, and cluster levels. Familiarity with Java/Go application environments and basic troubleshooting methods.
- Practical AWS production environment experience with core services (IAM, VPC, EC2, EKS, ALB/NLB, Route 53, S3, CloudWatch). Understanding of multi-AZ, high availability, network isolation, least privilege, auto-scaling, capacity planning, and cost governance principles. Capable of infrastructure design, risk assessment, and incident troubleshooting.
- Strong IaC, CI/CD, and automation skills with experience in Terraform/Terragrunt, GitHub Actions, Ansible, etc. Proficient in scripting (Bash/Python) to enhance infrastructure and application delivery reliability.
- Production observability and stability management experience with technologies like Prometheus, Grafana, ELK/Loki, OpenTelemetry. Understanding of Metrics, Logs, Tracing, SLI/SLO, alert management, capacity planning, and disaster recovery mechanisms.
- Familiarity with MySQL, Redis, Kafka operations and security practices (AWS IAM, Kubernetes RBAC, network/container security, Secrets/KMS, vulnerability management, security audits). Ability to identify and address production security risks.
- Web3/Crypto industry experience preferred, with understanding of blockchain fundamentals and digital asset trading security considerations.
Preferred Qualifications
- Experience in digital asset exchanges, securities, financial trading, or payment systems with high-availability requirements.
- Hands-on experience with large-scale/high-traffic Kubernetes/EKS environments, AWS multi-account governance, or cross-Region disaster recovery.
- Advanced cloud-native, cloud security, or cost optimization expertise (Karpenter, Argo CD/Flux, OpenTelemetry, eBPF, FinOps).
Benefits
- Highly competitive compensation package
- Singapore work visa sponsorship
- Flat organizational structure with diverse, inclusive team culture
- Generous vacation/sick leave policy with excellent work-life balance
Apply directly: https://davionlabs.bamboohr.com/careers/75?source=aWQ9MzM%3D