Live opening · Posted 12 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Company Description Cloudstrats is an Indian exponential technology product company that builds Artificial Intelligence platforms powered by analytics and automation. Its solutions use computer vision, speech, and text technologies to address complex real-world challenges for government and enterprise customers. The company’s platforms apply Deep Learning, NLP, and ML to support sustainable development goals across sectors such as citizen services, public safety, education, health, and transport. Cloudstrats delivers marquee projects including command control centers, e-governance solutions, grievance redressal platforms, data hubs, fraud detection, and health systems. The organization is driven by the vision of Atmanirbhar Bharat and large-scale digital transformation in India and beyond.
Immediate Joiners preferred.
About the Role.
We're looking for a hands-on private-cloud engineer who can take an OpenStack-based platform from bare metal to a working, accepted, and supportable production environment — and then own it through its operational lifecycle.
This is a greenfield deployment covering compute, software-defined storage, GPU enablement, Kubernetes, and an AI/ML platform. You'll lead or contribute to technical delivery, automation, capacity growth, operational improvement, resilience testing, documentation, and training of the internal cloud team.
This is not a role for someone who has only consumed OpenStack through a dashboard. We need someone who has deployed it, broken it, recovered it, automated it, and documented it — using Kolla Ansible as the deployment tooling.
The role combines architecture input with hands-on engineering. You won't be expected to single-handedly deliver every specialist subsystem without support, but you must understand the full platform and be able to troubleshoot across its major layers.
What You will do.
Deployment & Architecture
Design and deploy a multi-node OpenStack private cloud using Kolla Ansible
Configure a highly available control plane and define service placement, failure domains, and network segmentation (management, API, provider, tenant, external)
Deploy and configure core services: Keystone, Nova, Placement, Glance, Neutron, Cinder, Horizon
Maintain version-controlled inventory, configuration, and secrets-management procedures
Produce architecture documents, capacity plans, and implementation runbooks
Storage
Design and deploy a distributed software-defined storage platform (Ceph preferred) as the shared substrate
Integrate block/image storage with Glance, Cinder, and Nova; stand up an S3-compatible object endpoint
Define replication/erasure-coding policies, failure domains, and capacity thresholds
Validate failure, degraded-mode, and recovery behaviour
Networking & Security
Design and configure management, API, provider, tenant, storage, and tunnel networks
Implement VLANs, overlays, routing, firewalling, and security groups per the organisation's addressing standards
Apply security-hardening baselines and maintain compliance evidence
Compute & Workloads
Build and maintain golden images and flavour catalogues (CPU, memory, NUMA, huge pages, GPU profiles)
Configure host aggregates, availability zones, and placement/scheduling policies
Restore and re-platform existing Windows and Linux workloads into the new environment
GPU Enablement
Enable GPU access end-to-end: firmware, BIOS, IOMMU, kernel parameters, drivers, device binding
Configure PCI passthrough or vGPU as required; configure Nova PCI aliases, traits, and scheduler filters
Expose GPU resources to Kubernetes worker nodes via appropriate device plugins
Kubernetes & AI/ML Platform
Deploy and operate a highly available Kubernetes cluster on the private cloud
Configure CNI, ingress, RBAC, CSI-backed persistent storage, and GPU-schedulable nodes
Deploy an approved AI/ML workspace platform (notebooks, training jobs, model serving) with per-project isolation and quotas
Identity & Enterprise Integration
Integrate with the organisation's enterprise directory; configure SSO/federation across OpenStack, Kubernetes, and the AI workspace
Implement role mapping, least-privilege access, and credential-rotation procedures
Monitoring, Logging & Audit
Implement monitoring/alerting across hosts, OpenStack services, storage, VMs, Kubernetes, and GPU utilisation
Configure centralised logging, audit collection, and alert escalation
Resilience, Backup & Recovery
Define and execute resilience tests (compute/controller/storage/network/Kubernetes/GPU/identity failure scenarios)
Validate backup/restore against agreed RTO/RPO targets
Produce test plans, execution evidence, and resilience reports
Restricted-Network Delivery
Build and maintain internal package repositories and private container registries
Prepare an offline bill of materials (packages, images, charts, drivers, firmware, licenses)
Establish controlled import, scanning, signing, and promotion procedures
Ensure deployment, patching, and recovery can operate without dependency on public internet access
Lifecycle Ownership (post go-live)
Own technical evolution: capacity expansion, tuning, upgrades, automation
Act as escalation point for platform issues; lead root-cause analysis
Maintain architecture documents, SOPs, and runbooks
Train the internal cloud engineering team to independent competency
Experience we Need.
Core
6+ years infrastructure/platform-engineering experience; 3+ years hands-on OpenStack
Demonstrable delivery of at least one production OpenStack environment from install through acceptance
Evidence of personal technical contribution, not just project oversight
OpenStack & Kolla Ansible
Strong working knowledge of Keystone, Nova, Placement, Glance, Neutron, Cinder, Horizon
Hands-on Kolla Ansible: inventory, globals config, secrets handling, bootstrap, prechecks, deployment, validation, upgrades
Ability to troubleshoot containerised OpenStack services and deployment failures
Storage
Hands-on experience with a distributed SDS platform (Ceph preferred)
Experience integrating storage with Glance/Cinder/Nova and Kubernetes CSI
Linux, Compute & Networking
Deep Linux systems engineering (kernel, drivers, systemd, networking, storage stacks)
Strong KVM/QEMU understanding
VLANs, overlays, routing, firewalling, DNS, NTP, load balancing
Kubernetes & AI Platforms
Production Kubernetes deployment and operation experience
GPU scheduling via vendor device plugins
AI/ML platform deployment (notebooks, training, model serving) strongly preferred
GPU Virtualisation
PCI/IOMMU passthrough, device binding, driver installation and validation in VMs/containers
Identity, Observability & Automation
Enterprise directory integration, SSO/RBAC, certificate and secrets management
Strong Python/shell scripting and Ansible or equivalent IaC experience
Restricted-Network Environments
Demonstrable experience delivering in air-gapped or restricted-network environments
Comfortable with formal change control, security review, and audit evidence
Desirable.
Ceph RBD, CephFS, RADOS Gateway
Octavia, Barbican, Ironic, Heat, Swift, Manila
Kubeflow, JupyterHub, MLflow, Ray, KServe, vLLM, or comparable AI/ML platforms
NVIDIA GPU operators or equivalent vendor tooling
Multi-site disaster recovery experience
CKA / CKS / Linux / OpenStack certifications (supplementary to delivery evidence, not a substitute
You will own.
High-level and low-level architecture · Capacity-sizing workbook · Kolla Ansible configuration · IP/VLAN/DNS/firewall matrix · Storage architecture and recovery procedures · GPU enablement report · Kubernetes and AI-platform architecture · Restricted-network BOM · Security baseline · Monitoring/logging catalogue · Backup/RTO/RPO design · Migration and cutover plan · Resilience test plan and evidence · As-built documentation · Runbooks and SOPs · Training materials
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.