← 发现更多职位
字节跳动
Regular

AI Infra Engineer - Large Model Inference Systems (Multimodal / LLM / VLM)

面议
美国 · San Jose · 经验要求见详情
R&DRegularA122047美国国际招聘TikTok

关于这个机会

About the Team We are dedicated to building the inference infrastructure for ultra-large-scale language models, vision-language models, and frontier multimodal AI systems. Our mission is to provide a robust, scalable, and high-performance foundation for distributed serving, heterogeneous scheduling, and low-latency inference at massive scale. You will work on some of the most challenging problems in large-model online serving, spanning traffic orchestration, throughput and latency optimization, kernel efficiency, and production reliability for next-generation AI systems. Responsibilities - What You'II Do - Build and evolve next-generation inference systems for large-scale online traffic, including global scheduling across heterogeneous compute resources, high-concurrency load balancing, and efficient batch formation - Optimize distributed inference for 200B+ models and complex multimodal models through TP, EP, DP, and related strategies to improve throughput and latency in production - Develop high-performance kernels for frontier model architectures such as MoE, emerging attention mechanisms, and multimodal fusion layers using CUDA, Triton, and related tools - Explore AI-driven infrastructure for inference systems, including AI Agents for kernel optimization, performance tuning, consistency validation, deployment pipelines, and intelligent operations

任职要求

Minimum Qualifications: - Bachelor's degree or above in Computer Science, Software Engineering, Artificial Intelligence, Mathematics, or related fields - 2+ years of experience in high-performance computing, distributed scheduling systems, or large-model inference engine development - Familiarity with large-model architectures and strong system design skills for complex, high-concurrency environments - Strong understanding of asynchronous scheduling, resource pooling, and load balancing in distributed microservice systems - Strong engineering skills in performance optimization and production system development Preferred Qualifications - Deep understanding of inference frameworks such as vLLM and SGLang, with hands-on experience in customization and production optimization - Familiarity with GPU microarchitecture and operator-level optimization using CUDA, Triton, Cutlass, or related tools - Experience with LLM inference optimization, such as PTQ, QAT, KV cache optimization, or PD disaggregation - Experience deploying and optimizing VLMs or multimodal models in production

官方来源与核验

字节跳动官方招聘 · 职位编号 7651899647627807029

最近核验:2026-09-16T12:31:41.068845+00:00

查看官方职位详情 ↗

内推申请说明

本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。