Live opening · Posted 4 days ago

Principal Software Engineer

Rafay · United States (Remote)
Linkedin Yes
You are 4 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 4 days ago
CompanyRafay
LocationUnited States (Remote)
Salary4 benefits
Work modeYes
SourceLinkedin
Listed4 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
12 min from Linkedin publishing this role to us finding it
10 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
59,401 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Principal Software Engineer (Platform)
We are looking for a Principal Software Engineer who can make significant contributions to the design and development of the backbone of our multi-tenant SaaS, GPU PaaS, and Virtualization Kubernetes Operations Platform for a multi-cloud environment. Rafay is at the forefront of Kubernetes technology and we offer unique opportunities to develop new technology and be part of a team that encourages positive change through outside-of-the-box thinking. We hold high expectations for ourselves and challenge team members to continually seek improvement. Rafay offers opportunities to work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. As the first of our kind, we are truly in a class of our own.
As a Principal Engineer, you will serve as a technical leader and hands-on architect, setting the technical direction for some of the most critical services on the platform and multiplying the impact of the engineering teams around you.
Responsibilities
Design and Implement core architectural components for some of the most critical platform services of a multi-tenant distributed SaaS, GPU PaaS, and Virtualization platform
Set the technical vision and long-term architectural direction for major platform areas, and drive alignment across engineering teams
Build highly modular and scalable components and services for the platform
Perform R&D, feasibility analysis on latest technologies and newer versions of frameworks and libraries on an ongoing basis
Assist operations and solutions teams with deployment and stability of production systems
Collaborate with other team members and stakeholders including product management, UI designers and QA
Participate in and lead code reviews and design reviews, and establish engineering best practices and standards across teams
Mentor and provide technical guidance to senior and junior engineers, raising the technical bar across the organization
Enable a customer obsessed environment where team members can relentlessly champion and advocate for our customers in representing their issues to engineering teams and be a change agent to develop innovative ways to resolve their issues
Engineer and implement control plane services for GPU task scheduling, intelligent placement, and lifecycle handling across multi-cloud, diverse hardware infrastructure
Build highly scalable inference serving systems that deliver strict latency and throughput targets for multi-tenant deployments while optimizing GPU usage, capacity planning, and auto-scaling
Architect solution capabilities for model registries, artifact delivery, versioning strategies, blue-green or canary deployments, and concurrent LoRA inference execution across diverse GPU infrastructure
Skills and Qualifications
12+ years of experience in building and delivering large enterprise applications for customers
Deep understanding of distributed systems fundamentals, high availability and scalability principles
Expert knowledge of one or more of the following programming languages Golang, Python
Experience developing Micro-services
Strong troubleshooting and debugging skills
Strong understanding of multi-tenant isolation: per-tenant quotas, noisy-neighbor mitigation, and data, workload, and network isolation boundaries
Hands-on experience developing services on a public cloud platform (e.g., AWS, Azure, GCP)
Practical knowledge of networking protocols (TCP/IP, HTTP) and standard network architectures
Experience with containers and orchestration technologies like Kubernetes
Experience with GPU infrastructure, accelerated computing, or GPU scheduling/orchestration (e.g., NVIDIA GPU Operator, MIG, multi-tenant GPU sharing)
Experience with model registry and multi-cluster artifact distribution: model versioning, promotion across environments, and efficient replication of large artifacts to remote clusters is a plus
Understanding of LLM inference performance characteristics: continuous batching, KV cache management, prefill vs. decode behavior, quantization, and tensor/pipeline parallelism is a plus
Understanding of topology-aware scheduling and high-performance interconnects: NVLink, NCCL, RDMA/InfiniBand, GPUDirect, and their impact on placement decisions is a plus
Experience working and delivering features/enhancements/critical fixes for customer found issues
Demonstrated ability to influence technical direction and drive consensus across multiple teams and stakeholders
WHY JOIN RAFAY
Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App