About this role
Work on the inference engine layer and benchmarking infrastructure that runs models in production and partner environments; port models (e.g., LFM2) across runtimes and verify correctness. Design and build benchmark suites covering inference performance and model quality, run partner verifications, and maintain/extend inference engines built on llama.cpp, ONNX, and MLX.
Skills for this role
llama.cppONNX RuntimeMLXBenchmark designBenchmarking pipelinesBenchmarking methodologyModel portingModel verificationNumerical correctness verificationInference engine developmentC++PythonQuantizationDecoding strategiesMemory layoutEdge inference