Genesis Logo

Genesis

Inference

Posted 21 Days Ago
Be an Early Applicant
Hybrid
London, Greater London, England, GBR
Senior level
Hybrid
London, Greater London, England, GBR
Senior level
Design and implement low-latency, high-throughput inference pipelines for on-device and GPU-cluster deployment. Optimize kernels, batching, quantization, memory and scheduling. Build distributed serving systems, integrate low-level CUDA/Triton code with high-level frameworks, and develop monitoring/debugging tools to ensure reliability and determinism.
The summary above was generated by AI
What You’ll Do
  • Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics

  • Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization

  • Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks

  • Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)

  • Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)

  • Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)

  • Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling

  • Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments

  • System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness

Similar Jobs

15 Days Ago
Remote or Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Fintech • Software • Financial Services • Cryptocurrency
Optimize large-model inference at scale: improve throughput, reduce latency and cost per token, build benchmarking harnesses, tune parallelism and quantization strategies, implement load-balancing in routing, and evaluate custom kernels and emerging inference hardware.
Top Skills: AsicsB300 NodesC++Continuous BatchingCudaCuda GraphsCustom Triton KernelsFpgasGoKv Cache ManagementMoe ParallelismNsight ComputeNsight SystemsPagedattentionPipeline ParallelismPythonPytorch ProfilerQuantizationRustSglangSpeculative DecodingTensor ParallelismTorch.CompileTritonVllm
13 Minutes Ago
In-Office
Senior level
Senior level
Artificial Intelligence • Legal Tech • Software
Lead GTM team revenue performance, quota attainment, pipeline management, forecasting, deal inspection, and execution standards. Coach and develop team members, identify conversion and performance gaps, and use analytics and AI-driven insights to improve productivity, deal velocity, win rates, and forecast accuracy. Partner with Marketing, Product, Customer Success, and Revenue Operations to strengthen funnel performance, refine strategy, and improve GTM playbooks. Requires B2B SaaS experience and English-French fluency.
Top Skills: AIB2B Saas
6 Hours Ago
Hybrid
London, Greater London, England, GBR
Mid level
Mid level
eCommerce • Information Technology • Marketing Tech • Software
Build and operate end-to-end product features across a Go backend on Google Cloud Run and React frontend. Develop LLM-powered analysis and recommendations, evaluate nondeterministic outputs, design prompts, and create independent connector services for external data sources. The role includes production ownership, system design, customer feedback, experimentation, code reviews, and collaboration with senior engineers.
Top Skills: Akeneo PimGoGoogle Cloud RunLarge Language ModelsPrompt EngineeringReact

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account