About this role
Together AI is seeking a Staff Engineer to operate, scale, and optimize multi-petabyte storage systems designed for large-scale AI training and inference workloads. You will lead the technical strategy for storage infrastructure, managing high-performance parallel filesystems and object stores to support extreme throughput and low-latency data paths. Responsibilities include architecting storage roadmaps, developing Kubernetes-native storage operators, and implementing intelligent caching architectures to ensure performance at GPU scale. You will also focus on multi-tenancy, quota enforcement, and cost optimization through automated tiering. Requirements: 8+ years of experience in storage engineering with distributed systems at multi-petabyte scale; proven track record with GPU/HPC clusters; deep expertise in Kubernetes and cloud-native storage; strong proficiency in Go and Python; and experience with parallel filesystems like Ceph, WekaFS, or Lustre. You should have a history of technical leadership in improving system performance and reliability. Preferred skills include GPU Direct Storage (GDS), NVMe-oF, and storage benchmarking tools like fio and iperf3. A BS/MS in Computer Science or equivalent is required.