Research Engineer, LLM/VLM Inference Optimization (Kernel & Compiler) - Seed Infra
关于这个机会
About the Team The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models. Responsibilities - Design, implement, and optimize high-performance GPU kernels for large-scale LLM/VLM inference workloads, including attention, GEMM, and other compute- and memory-intensive operators. - Develop and tune inference kernels in CUDA and Triton, and drive end-to-end performance optimization of production inference systems at scale. - Conduct in-depth performance analysis and profiling to identify bottlenecks across the inference stack, from kernel level to serving level. - Collaborate with research and infrastructure teams to land kernel- and compiler-level optimizations in production inference systems.
任职要求
Minimum Qualifications: - Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field. - Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming. - Hands-on experience in LLM/VLM inference optimization with demonstrated impact on latency, throughput, or serving cost. - Hands-on experience writing and optimizing GPU kernels in CUDA and/or Triton. - Deep understanding of GPU architecture (memory hierarchy, occupancy, instruction throughput) with solid optimization experience. Preferred Qualifications: - Experience with ML compiler internals (e.g., Triton, MLIR, LLVM). - Contributions to related open-source projects (e.g., Triton, vLLM, SGLang, FlashAttention, CUTLASS). - Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).
内推申请说明
本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。