Job Opportunity Posted today

VMware Operations Lead

Kotak Mahindra Bank
Mumbai

Job Description

Role : VMware Operations Lead

Location: Mumbai/ Navi Mumbai

Years of experience: 9+


ROLE PURPOSE:

Lead the VMware virtualization and private-cloud operations function, ensuring availability, performance, capacity, security, compliance, lifecycle health and service continuity across Production, Disaster Recovery, UAT and other managed environments. The role combines hands-on technical depth with people leadership, operational governance and platform modernization. Critical success profile: A hands-on leader who can independently diagnose complex cross-domain issues across compute, storage, networking, backup, operating systems and platform integrations, while guiding the team through restoration and permanent resolution.


KEY RESPONSIBILITIES:

VMware Platform Operations:

  • Own day-to-day operations of enterprise VMware environments comprising ESXi hosts, vCenter Server, clusters and virtual machines.
  • Ensure platform availability, reliability, performance, capacity, resiliency and compliance across Production, DR and non-production environments.
  • Lead installation, configuration, hardening, patching, upgrades, lifecycle management, technology refresh and decommissioning activities.
  • Perform proactive health checks, trend analysis, capacity reviews and performance optimization; translate findings into prioritized remediation plans.
  • Act as the senior technical escalation point for critical incidents and coordinate timely service restoration.

Troubleshooting & Problem Management :

  • Demonstrate strong hands-on troubleshooting across VMware compute, memory, storage, networking, guest operating systems, backup and hardware layers.
  • Use logs, metrics, events and dependency analysis to isolate root cause, validate hypotheses and implement sustainable fixes.
  • Lead major-incident technical bridges, root-cause analysis, corrective and preventive actions, and knowledge transfer.
  • Troubleshoot complex cross-domain issues without depending solely on OEM or vendor support.

VMware Cloud Foundation (VCF):

  • Operate and support VMware Cloud Foundation, including SDDC Manager, vSphere, vCenter, NSX, vSAN and relevant VMware Aria components.
  • Manage VCF workload domains, platform health, certificate/password lifecycle, patching, upgrades and configuration compliance.
  • Contribute to private-cloud standardization, modernization and operational-readiness initiatives.

Virtualization Networking:

  • Operate and troubleshoot vSphere Distributed Switches (vDS), port groups, uplinks, teaming, VLANs and virtual network policies.
  • Apply working knowledge of NSX concepts including segments, gateways, routing, distributed firewall, micro-segmentation, overlay/underlay networking and service connectivity.
  • Understand and support NSX VPC concepts, tenant isolation and policy-driven networking where deployed.
  • Coordinate with network and security teams for end-to-end connectivity, firewall and performance troubleshooting.

Automation & Platform Engineering:

  • Identify repetitive operational activities and convert them into controlled, reusable automation.
  • Use or guide automation through PowerCLI, PowerShell, Python, REST APIs, VMware Aria Automation, Aria Orchestrator and relevant Infrastructure-as-Code practices.
  • Promote automated provisioning, Day-2 operations, compliance validation, reporting and evidence generation.
  • Apply version control, peer review, testing, rollback and documentation practices to production automation.

Storage, Backup & Disaster Recovery:

  • Work with storage teams on SAN, NAS, vSAN, datastore performance, multipathing, capacity and storage-policy issues.
  • Operate or support integration with enterprise VM backup and recovery tool like Veeam or equivale
  • Participate in recovery validation, DR drills, runbook maintenance and closure of identified gaps.

Team Leadership & Service Management:

  • Lead, mentor and develop VMware administrators and engineers; allocate work and strengthen technical depth within the team.
  • Govern Incident, Problem, Change, Request, Capacity, Availability and Configuration Management processes in line with defined SLAs/OLAs.
  • Coordinate with application, network, security, storage, database, backup, monitoring, hardware and project teams.
  • Manage OEM/vendor escalations and ensure clear ownership, evidence, follow-up and closure.

Documentation, Risk & Compliance:

  • Create, review and maintain SOPs, runbooks, support matrices, architecture diagrams, build standards, troubleshooting guides and knowledge articles.
  • Ensure changes, approvals, technical evidence and operational records remain traceable and audit-ready.
  • Support security hardening, vulnerability remediation, access review, audit requests and policy compliance.
  • Identify platform risks, lifecycle constraints, capacity concerns and single points of failure; drive treatment plans to closure.


MANDATORY QUALIFICATIONS & EXPERIENCE:

  • 10+ years of hands-on experience working in VMware-based enterprise virtualization environments.
  • Minimum 4 years of experience managing or technically leading a VMware administration/operations team.
  • Strong and demonstrable knowledge of VMware vSphere fundamentals, including ESXi, vCenter, HA, DRS, vMotion, Storage vMotion, resource management, permissions and lifecycle operations.
  • Practical experience in large VMware environments across multiple deployment models and criticality tiers.
  • End-to-end managed-environment lifecycle experience covering installation, configuration, hardening, monitoring, support, troubleshooting, upgrades and decommissioning.
  • Hands-on working knowledge of vSphere networking, vDS and NSX; understanding of NSX VPC concepts is required or expected to be developed quickly.
  • Working experience with VMware Cloud Foundation and understanding of SDDC Manager and core VCF components. VMware Operations LeadThis is a Confidential document. VMware Operations Lead.
  • Practical knowledge of x86 rack/blade server hardware and associated firmware, BIOS, management-controller and hardware diagnostic concepts.
  • Basic knowledge of enterprise storage, Windows and Linux operating systems.
  • Good understanding of IT service delivery, Incidents, Problems, Changes, Requests and SLAs.
  • Ability to communicate clearly with technical teams, management, vendors and business stakeholders.


CORE TECHNICAL COMPETENCIES:


Working knowledge:

  • Server Hardware - x86 blade/rack servers, firmware, diagnostics and vendor coordination
  • OS Platforms - Windows and Linux administration and troubleshooting fundamentals
  • Backup & Recovery - VM backup integration, recovery validation and failure troubleshooting

Advanced / hands-on:

  • VMware vSphere - Fundamentals, administration, performance, lifecycle and troubleshooting
  • Troubleshooting - Cross-domain fault isolation, root cause analysis and resolution leadership

Strong working knowledge/ Strong aptitude:

  • VCF / SDDC Manager - Workload domains, lifecycle, health and component integration
  • vDS , NSX / NSX VPC - Virtual networking, segmentation, routing, firewall and tenant isolation concepts
  • VMware Aria Suite - Operations, Automation, Orchestrator and Logs / log analytics
  • Automation - PowerCLI, scripting, APIs, orchestration and controlled automation practices


PREFERRED / ADDED ADVANTAGE:

• Understanding of Kubernetes fundamentals; exposure to VMware Kubernetes Service (VKS), Tanzu or another enterprise container platform is preferred.

• Experience with Infrastructure as Code, configuration management or automation frameworks such as Terraform and Ansible.

• Exposure to hybrid-cloud, private-cloud operating models, DevOps practices and CI/CD concepts.

• Experience with HCX, SRM or equivalent migration and disaster-recovery technologies.

• Experience in BFSI, financial services or another highly regulated enterprise environment.

• Relevant VMware certifications such as VCP-DCV, VCP-VCF, VCAP or NSX certification.


LEADERSHIP & BEHAVIOURAL COMPETENCIES:

• Provides calm, structured and accountable leadership during critical incidents.

• Balances operational stability with modernization and automation objectives.

• Challenges assumptions through evidence and communicates risk in business-relevant language.

• Develops team capability through coaching, delegation, reviews and knowledge sharing.

• Collaborates effectively across infrastructure, application, security, audit and vendor teams.

• Demonstrates ownership, integrity, customer focus and disciplined follow-through.

About this job listing
This job opportunity is provided through our external job listing network. MyJobAlerts helps you discover job opportunities and redirects you to the original listing to apply.