Live opening · Posted 5 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Vision Solution Architect
Computer Vision, VLMs & Multimodal GenAI Architecture
Experience Level: 8-12 years
Employment Type: Full-Time
Location: Pune, Chennai, Hyderabad, Kolkata, Delhi, Bangalore
About the Role
We are looking for a Vision Solution Architect to design and lead end-to-end computer vision and multimodal AI solutions — spanning classical vision models, CLIP-style embedding models, Vision-Language Models (VLMs), and Vision RAG architectures. The ideal candidate can architect solutions across cloud hyperscalers and edge deployments, guiding fine-tuning, integration, and productionization of vision and multimodal GenAI systems for real-world business use cases.
Key Responsibilities
Architect end-to-end vision and multimodal GenAI solutions, from use-case discovery and model selection through deployment and monitoring.
Design solutions leveraging CLIP-style embedding models, VLMs (e.g., LLaVA, Gemini Vision, GPT-4V/Claude Vision, Qwen-VL), and traditional CV models (detection, segmentation, OCR).
Design and implement Vision RAG pipelines — multimodal embedding generation, vector indexing/retrieval, and grounding of VLM outputs against visual and textual knowledge bases.
Lead fine-tuning and domain adaptation of vision and vision-language models (LoRA/QLoRA, adapter-based tuning, contrastive fine-tuning) for domain-specific accuracy.
Architect solutions across hyperscaler vision/AI services — AWS (Rekognition, Bedrock multimodal models, SageMaker), GCP (Vertex AI Vision, Gemini multimodal APIs), and Azure (AI Vision, Azure AI Foundry) — selecting the right managed vs. self-hosted approach.
Design edge deployment architectures for vision models on constrained devices (NVIDIA Jetson, edge TPUs, mobile/embedded accelerators), including model compression and quantization.
Select and optimize model formats, runtimes, and accelerators (ONNX, TensorRT, OpenVINO, CoreML) for target hardware and latency/throughput requirements.
Define evaluation frameworks for vision/multimodal model quality — embedding retrieval accuracy, grounding fidelity, hallucination, and task-specific benchmarks.
Collaborate with data science, MLOps, and client-facing teams to translate business requirements into scalable, cost-efficient vision architectures.
Document reference architectures, best practices, and reusable accelerators for vision and multimodal GenAI engagements.
Required Skills & Experience
Strong hands-on experience with CLIP and other contrastive embedding models, and Vision-Language Models (VLMs) for tasks such as captioning, visual QA, and grounding.
Practical experience designing Vision RAG systems, including multimodal embeddings, vector databases, and retrieval-augmented generation patterns.
Experience fine-tuning and adapting vision/VLM models using parameter-efficient techniques (LoRA/QLoRA) and classical CV model training/transfer learning.
Working knowledge of hyperscaler AI/vision services across at least two of AWS, GCP, and Azure, and their trade-offs for vision workloads.
Familiarity with edge AI deployment — hardware accelerators (NVIDIA Jetson, Coral/edge TPU), model compression, quantization, and runtime optimization.
Proficiency in Python and vision/ML frameworks (PyTorch, Hugging Face Transformers, OpenCV) and model format/runtime tooling (ONNX, TensorRT, OpenVINO).
Understanding of containerization and orchestration (Docker, Kubernetes) for scalable vision model serving.
Strong architectural and communication skills, with the ability to translate business needs into technical vision/GenAI solution designs.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.