Live opening · Posted 11 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Title: Data Solution Architect – Air-Gapped On-Premise (Data Warehouse / Data Lake / Lakehouse)
Company: Cloudstrats
Location: [Delhi] · [On-site]
Employment Type: Full-time
Experience: [10+] years
About the Role
We're looking for a seasoned Data Solution Architect to design and lead secure, fully on-premise data platforms in air-gapped (disconnected) environments. You'll architect ETL, data warehouse, data lake and lakehouse, advanced analytics and dashboard solutions that run with no internet or cloud dependency. The role suits someone who enjoys building strong platforms under strict security, compliance and change-control constraints, for clients in sectors such as government, defense, BFSI, energy and critical infrastructure.
Key Responsibilities
Define end-to-end on-premise data architecture and roadmaps across ingestion, storage, processing, governance, analytics and visualization layers, all within air-gapped networks
Design enterprise data warehouses (Kimball, Inmon, Data Vault 2.0) on self-hosted platforms such as PostgreSQL, Greenplum, ClickHouse, Oracle, SQL Server, Teradata or Vertica
Architect on-premise data lakes and lakehouses using object storage (MinIO, Ceph, HDFS), open table formats (Apache Iceberg, Delta Lake, Hudi), catalogs (Hive Metastore, Nessie) and query/processing engines (Apache Spark, Trino/Presto, Dremio, Flink)
Design batch and streaming pipelines using open-source and self-hosted tools such as Talend, Pentaho/Apache Hop, Apache NiFi, Airflow, dbt Core and Kafka
Plan secure data ingestion across network boundaries, including cross-domain solutions, data diodes, controlled media transfer, and file validation and scanning workflows
Set up offline software supply chains: internal package and container mirrors (Nexus, Artifactory, Harbor), offline Python/Maven/OS repositories, and controlled patch and upgrade procedures
Deploy platforms on bare metal, VMware or on-premise Kubernetes (OpenShift disconnected, RKE2, Rancher), with infrastructure sizing, capacity planning and HA/DR design
Enable on-premise advanced analytics and AI/ML with JupyterHub, MLflow, Kubeflow and Spark MLlib, and optionally self-hosted LLMs (vLLM, Ollama) with local vector databases
Guide BI and dashboard strategy with self-hosted tools such as Qlik Sense Enterprise on Windows, Power BI Report Server, Tableau Server and Apache Superset, including semantic layers and self-service analytics
Establish security and governance frameworks: LDAP/AD/FreeIPA, Kerberos, Keycloak, Apache Ranger, HashiCorp Vault, encryption at rest and in transit, data masking, audit logging, and data catalog/lineage (Apache Atlas, OpenMetadata, DataHub)
Set up on-premise observability and operations using Prometheus, Grafana, the ELK/OpenSearch stack and alerting, plus backup and recovery strategies
Align architectures with security and compliance requirements (e.g., ISO 27001, CERT-In guidelines, RBI/SEBI norms, DPDP Act, sector-specific security standards) and support security audits and accreditation
Lead modernization of legacy on-premise warehouses to lakehouse architectures within disconnected environments
Engage in pre-sales and solutioning: RFP responses, bill of materials, hardware/licensing estimates, PoCs and technical presentations
Mentor architects, data engineers and BI developers, and define architecture standards, runbooks and operational procedures
Required Skills & Qualifications
[10+] years in data engineering, data warehousing or BI, with at least [4+] years in a solution or data architect role
Proven experience designing and delivering on-premise data platforms, including in air-gapped, restricted or highly regulated environments
Deep knowledge of data modeling (dimensional, Data Vault, normalized) and modern data architecture patterns (lakehouse, data mesh, data fabric)
Hands-on expertise with the open-source data ecosystem: Spark, Trino, Kafka, NiFi, Airflow, Iceberg/Delta, MinIO/HDFS
Strong experience with on-premise ETL tools such as Talend, Pentaho or Apache Hop
Strong Linux administration knowledge (RHEL/Rocky/Ubuntu), networking fundamentals and security hardening
Experience with containers and on-premise Kubernetes, including offline image and package management
Advanced SQL, plus proficiency in Python and/or Scala/Java
Working knowledge of self-hosted BI and visualization tools
Excellent communication, documentation and stakeholder management skills
Bachelor's or Master's degree in Computer Science, IT, Engineering or a related field
Good to Have
Experience with Cloudera (CDP Private Cloud), Hadoop distributions or other commercial on-premise data platforms
Infrastructure as Code and automation in disconnected setups (Ansible, Terraform with local providers)
Hands-on experience with data diodes, cross-domain solutions or secure file-transfer gateways
Exposure to GPU infrastructure for on-premise AI/ML and LLM workloads
Domain experience in government, defense, BFSI, telecom, energy or healthcare
Certifications such as TOGAF, Red Hat (RHCE/OpenShift), CKA/CKS, Cloudera, CISSP or other security certifications
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.