Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Job Requirements
About the Role
The Machine Learning (ML) Engineer will be part of the Data & Analytics team, based in Mumbai or Bangalore . This role is focused on managing the training infrastructure for Large Language Models (LLMs), such as GPT-4 and BERT. The ideal candidate will have deep expertise in machine learning, GPU optimization, and distributed computing. You will work closely with data scientists and ML engineers to ensure efficient, scalable, and reliable training environments for large-scale AI models.
Key Responsibilities
Primary Responsibilities
Manage and optimize infrastructure for training large-scale machine learning models, particularly LLMs.
Leverage deep learning frameworks such as TensorFlow, PyTorch, or Keras for model training.
Maximize GPU utilization and efficiency through a strong understanding of computer architecture.
Utilize cloud platforms (AWS, Azure, GCP) to manage and scale training environments.
Apply containerization and orchestration tools like Docker and Kubernetes for deployment and resource management.
Implement parallel and distributed computing principles to support large model training.
Apply MLOps practices to manage the end-to-end machine learning lifecycle.
Secondary Responsibilities
Work on infrastructure for deep learning projects involving models like BERT and Transformers.
Optimize GPU resources both on-premise and in the cloud.
Troubleshoot hardware and performance issues during model training.
Collaborate with data scientists to understand infrastructure needs and deliver efficient solutions.
What We Are Looking For
Education
Graduation in BSC or BCA or B.Tech
Experience
2+ years of relevant experience in machine learning engineering, with a focus on infrastructure and model training.
Experience in managing large-scale ML projects, preferably involving LLMs.
Skills and Attributes
Expertise in deep learning frameworks and GPU optimization.
Strong knowledge of distributed computing and big data technologies (e.g., Hadoop, Spark).
Familiarity with MLOps tools and practices.
Excellent problem-solving and collaboration skills.
Key Success Metrics
Timely and error-free delivery of infrastructure solutions.
Effective identification and resolution of training infrastructure issues.
Technical leadership in ML infrastructure projects.
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.