Job Opportunity Posted today

Senior DevOps Engineer

Cogrion
Athiyur

Job Description

Domain: Cloud | Kubernetes | Data Platforms | Platform Engineering

Cogrion is building an Autonomous Data & AI Infrastructure Platform for modern enterprises.

We’re looking for a Senior DevOps Engineer who can design, automate, secure, and operate large-scale cloud-native platforms across Kubernetes, data infrastructure, and AI workloads.


What You’ll Work On

  • Design and operate production-grade Kubernetes platforms
  • Build and manage infrastructure using Terraform / OpenTofu and Helm
  • Implement GitOps and CI/CD using ArgoCD, GitHub Actions, GitLab CI, or similar
  • Manage workloads across AWS, Azure, GCP, and Alibaba Cloud
  • Design autoscaling, capacity management, and cost optimization strategies
  • Operate platforms running Spark, Trino, Airflow, Jupyter, MLflow, Kafka, and related services
  • Build observability using Prometheus, Grafana, Loki, OpenTelemetry, and alerting systems
  • Improve platform reliability, availability, security, and disaster recovery
  • Implement IAM, workload identity, secrets management, network security, and RBAC
  • Troubleshoot complex Kubernetes, networking, storage, and distributed-system issues
  • Automate platform deployment, upgrades, patching, backup, and recovery

What We’re Looking For

  • 7–10 years of DevOps / SRE / Platform Engineering experience
  • Strong Kubernetes and Docker expertise
  • Strong experience with AWS and at least one additional cloud platform
  • Terraform / OpenTofu, Helm, and GitOps
  • CI/CD design and automation
  • Linux, networking, DNS, ingress, load balancing, and cloud networking
  • IAM, security, secrets, certificates, and workload identity
  • Strong scripting skills in Python, Bash, or similar
  • Production troubleshooting and incident-management experience
  • Strong understanding of scalability, reliability, and infrastructure cost optimization


Bonus

Experience with:

  • EKS / AKS / GKE / Alibaba ACK
  • Karpenter and Kubernetes autoscaling
  • Spark and distributed data workloads on Kubernetes
  • Trino, Airflow, JupyterHub, MLflow, Kafka
  • Keycloak, OIDC, IRSA / workload identity
  • Prometheus, Grafana, Loki, OpenTelemetry
  • ArgoCD
  • Multi-tenant SaaS or enterprise platforms
  • SOC 2 / ISO 27001 / cloud security practices

What Matters to Us

We’re looking for someone who can go beyond infrastructure operations and think like a platform engineer and product owner.

You should be comfortable taking ownership from:

Architecture → Automation → Security → Deployment → Observability → Production Reliability

At Cogrion, you’ll work at the intersection of:

Cloud Infrastructure × Kubernetes × Data Platforms × AI Infrastructure

About this job listing
This job opportunity is provided through our external job listing network. MyJobAlerts helps you discover job opportunities and redirects you to the original listing to apply.