About this role
Beam seeks an Applied AI Research Engineer to lead hands-on research into LLM inference, reducing cost per token and latency for customer workloads on its serverless, GPU-backed inference platform. The role is listed in New York, NY, or San Francisco, CA, with a remote option within the US.
Responsibilities
- Optimize low-level inference using techniques such as speculative decoding, quantization, KV-cache optimization, and memory management.
- Work directly with customers to improve production workloads, then turn those findings into platform improvements.
- Independently identify high-impact research opportunities and help guide the future of Beam’s inference platform.
Requirements
- At least 3 years of experience, as listed in the posting, and a systems or research background in LLM inference.
- Deep understanding of LLM serving, from kernels through scheduling, and a record of shipping products or research used in production-like academic or industry settings.
- Willingness to collaborate closely with customers and enthusiasm for developer tools, cloud-native technologies, and open-source software. The listing also names PyTorch, reinforcement learning, and GPU programming as skills. Applicants must be US citizens or hold a US visa.
Compensation and benefits: Listed salary is $140K–$200K, with 0.25%–1.00% equity. Beam offers health, dental, and vision benefits with 90% coverage for employees and 50% for dependents, plus a fitness stipend, learning budget, and opportunities to participate in cloud-native community events.