Original job description
About the Role
PolarGrid is building the infrastructure layer for real-time AI inference. We're looking for an Inference Optimization Engineer to squeeze every bit of performance out of our stack. You'll work directly on the systems that serve inference to our customers, making them faster, cheaper, and more efficient.
This is a deep technical, systems-focused role. You'll own latency, throughput, and cost per token as real metrics you're responsible for improving. You'll build the repeatable benchmarking and optimization process that takes new models and hardware from an initial baseline to a validated production configuration.
What You'll Do
Profile and optimize inference pipelines end to end using representative customer workloads, from request handling and scheduling through distributed GPU execution
Tune serving frameworks such as vLLM, TensorRT-LLM, and SGLang for specific latency, throughput, and cost targets
Build automated benchmarking and performance regression tooling across models, frameworks, precisions, hardware, and workload profiles
Characterize customer workloads and translate TTFT, ITL, concurrency, and context-length requirements into deployment configurations
Implement and evaluate quantization strategies across model families, measuring both performance gains and model-quality regressions
Work with hardware teams to match model configurations and parallelism strategies to GPU topology, NVLink, and interconnect bandwidth
Benchmark new hardware such as RTX Pro 6000s and B300s, identifying the best engine, precision, parallelism, and deployment configuration for each workload
Bring new model architectures into production, including checkpoint conversion, framework support, distributed configuration, and correctness validation
Contribute to continuous batching, speculative decoding, KV-cache optimization, prefill/decode disaggregation, and request-scheduling work
Read, debug, and modify inference framework internals when configuration-level tuning is not enough
Work with the platform team to canary performance improvements, measure them under production traffic, and turn successful configurations into repeatable deployment recipes
What We're Looking For
Strong GPU systems fundamentals, with the ability to work across Python, C++, CUDA, or Triton when optimization requires going below framework configuration
Hands-on experience with at least one major inference serving framework such as vLLM, TGI, TensorRT-LLM, or SGLang
Deep understanding of transformer architecture and where inference bottlenecks actually live
Ability to read, debug, and modify inference framework internals rather than treating them as black boxes
Experience building benchmarking, load-generation, or performance-regression infrastructure
Comfortable profiling with Nsight Systems, Nsight Compute, PyTorch Profiler, or similar tools
Experience with quantization and precision tradeoffs in production, including validating numerical correctness and model quality
Experience optimizing multi-GPU or multi-node inference across high-speed interconnects
Understanding of distributed inference, NCCL, GPU topology, and communication bottlenecks
You care about numbers: TTFT, ITL, P95/P99 latency, throughput, GPU utilization, and tokens/sec/dollar
Bonus Points
Experience writing custom CUDA or Triton kernels
Familiarity with speculative decoding, MoE routing optimizations, or prefill/decode disaggregation
Experience with inference request routing, scheduling, or admission control
Experience upstreaming performance improvements to vLLM, SGLang, TensorRT-LLM, or related projects
Open-source contributions to inference or ML systems projects
Why PolarGrid
You'll work on real hardware at scale, not toy benchmarks. The performance improvements you ship go directly to customers and directly affect our unit economics. Small team, real ownership.