Cloud Platform Engineer

Bengaluru, Karnataka, India | Engineering | Full-time

Apply

Cloud Platform Engineer (AWS / EKS / Terraform)

**Team:** Infrastructure · Reports to Head of Infrastructure

About the role
You will build and operate TookiTaki's cloud platform: AWS EKS clusters managed as code, GitOps-driven delivery with the Argo suite, operator-managed data services running on Kubernetes, and a Terraform-based self-service layer that lets application teams request infrastructure through reviewed YAML instead of tickets. The job is platform engineering, not click-ops — everything ships through version control, automated pipelines, and policy gates.

Requirements
Education
- Required: Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent practical experience.
- Preferred: Relevant certifications over a Master's — CKA (Certified Kubernetes Administrator), HashiCorp Terraform Associate, AWS Solutions Architect Associate or higher.

Experience
- 3–5+ years in cloud, platform, or DevOps engineering.
- Proven track record running production Kubernetes on a managed cloud (EKS strongly preferred), including stateful workloads.
- Experience operating infrastructure entirely through Infrastructure as Code — no console-driven change management.

Technical expertise
- **Kubernetes / EKS:** cluster lifecycle and upgrades, autoscaling (Karpenter, KEDA, VPA), ingress and load balancing, IRSA/Pod Identity, core add-ons (cert-manager, external-dns, CoreDNS, node-local-dns).
- **GitOps / Argo:** ArgoCD for platform and application delivery (app-of-apps, Helm chart authoring, sync and rollback strategies), Argo Rollouts for progressive delivery, Argo Workflows/Events for automation.
- **Databases and stateful services on Kubernetes:** deploying and operating databases via Kubernetes operators — PostgreSQL and MySQL (Percona operators), ScyllaDB, Elasticsearch, Valkey/Redis, Kafka (Strimzi) — plus AWS RDS/Aurora where managed services fit better; backup/restore, upgrades, and capacity management for stateful workloads.
- **Terraform:** authoring reusable, versioned, tested modules (not just consuming them) — variable/output interface design, `terraform test`, semantic versioning, remote state on S3.
- **CI/CD & automation:** pipeline design in GitHub Actions or GitLab CI; plan/apply automation with Atlantis; policy-as-code gates (OPA/Conftest, checkov, tflint); pre-merge validation and drift detection.
- **AWS core services:** VPC and network design, Route 53, IAM (least-privilege roles and policies), S3, RDS/Aurora.
- **Observability:** Datadog, OpenTelemetry (collector and kube-stack), Grafana; alerting hygiene with low false-positive rates; log and event pipelines.
- **Programming:** solid scripting/tooling ability in Python or Go — enough to build renderers, validators, and pipeline tooling, not just glue scripts.
- **Nice to have:** data platform tooling (Airflow, Spark on Kubernetes, Temporal, StarRocks), identity and SSO (Keycloak, Dex, oauth2-proxy), load testing with k6, FinOps practices (tagging standards, cost allocation, rightsizing).

Soft skills
- Strong problem-solving and analytical abilities; comfortable debugging across the stack (DNS → LB → cluster → workload → database).
- Clear written communication — design docs, runbooks, and PR descriptions are first-class deliverables here.
- Collaborative mindset: infrastructure changes ship through peer review, and platform decisions are made with (not for) application teams.

Key competencies
- **Platform thinking:** design paved roads that make the secure, cost-efficient path the easy path for application teams.
- **Automation-first:** if a task is done twice manually, the third time is a pipeline.
- **Ownership:** own services end to end — provisioning, upgrades, incidents, cost, and documentation.
- **Cost awareness:** treat cloud spend as an engineering metric; tag, measure, and optimize continuously.
- **Adaptability:** comfortable in a fast-moving environment where the platform itself is under active development.

Success metrics (first 6–12 months)
- Application teams provision standard infrastructure through the self-service platform with no manual Terraform written by requesters.
- EKS cluster and node-group upgrades executed as routine, zero-downtime operations via blue/green rollout.
- Operator-managed data services (PostgreSQL, Kafka, Elasticsearch, etc.) run with tested backup/restore and rehearsed upgrade procedures.
- 100% of infrastructure changes delivered through reviewed, policy-gated pipelines — zero out-of-band console changes.
- Deployment lead time for platform changes reduced measurably (target: same-day merge-to-production for routine changes).
- Cost visibility established via tagging and dashboards, with identified savings executed (rightsizing, autoscaling, storage tiering).
- Actionable alerting: on-call pages correspond to real incidents; false-positive alerts trend toward zero.
- Mean time to recovery for platform incidents under 30 minutes.

Benefits
- **Competitive salary** aligned with industry standards and experience.
- **Professional development:** certification support (CKA, AWS, Terraform) and training across cloud, platform, and data engineering.
- **Comprehensive benefits:** health insurance and flexible working options.
- **Growth opportunities:** career progression within TookiTaki's expanding infrastructure and platform organization.