About this role
Fireworks seeks an AI Field Engineer for its Enterprise track to help large organizations and digital-native companies turn generative AI projects into production systems. The role combines hands-on engineering, solution design, customer delivery, and executive-level communication. Engineers embed with customer teams, including on-site work, and own technical relationships from discovery through deployment.
Responsibilities
- Build end-to-end proofs of concept, MVPs, and production integrations within customers’ codebases and infrastructure. Architect inference foundations and size deployments for GenAI products.
- Load-test realistic traffic; establish latency, throughput, and cost baselines; and tune deployments against those targets. Deploy and validate model families using frameworks such as vLLM and SGLang, selecting appropriate configurations, quantization, and serving patterns.
- Advise on model selection, fine-tuning approaches including SFT, DPO, and RFT, and evaluation methodology. Build fine-tuning pipelines and production-quality evaluation frameworks while balancing quality and compute cost.
- Lead discovery conversations, align engineering and executive stakeholders, debug production issues, and translate recurring customer needs into product proposals, deployment patterns, documentation, and roadmap feedback.
Required qualifications
- 5+ years in a hands-on, customer-facing technical role, such as forward-deployed engineering, applied AI engineering, solutions architecture, ML engineering with field exposure, or technical founding. Must have built and shipped software in a customer’s production environment.
- Strong Python and production debugging skills; familiarity with Kubernetes and infrastructure engineering; working knowledge of LLM inference, model serving, and fine-tuning, including SFT. Experience with AWS, Azure, or GCP and deploying models on GPU infrastructure. Strong communication across technical and executive audiences.
Preferred qualifications
- 10+ years in technical field or engineering roles; experience with vLLM, SGLang, or TensorRT-LLM, hyperscaler AI platforms, agentic systems or tool-use chains, and taking GenAI prototypes to production scale. DPO or RFT experience is a strong plus.
The posting lists San Mateo, New York, and United States remote locations, labels the location type Hybrid, and expects customer-site work. Compensation is $200K–$260K in on-target earnings plus equity.