← 发现更多职位
字节跳动
Regular

Backend Engineer , AML Engine Orchestration

面议
新加坡 · Singapore · 经验要求见详情
R&DRegularA225574A新加坡国际招聘TikTok

关于这个机会

Team Introduction The mission of our AML team is to push next-generation machine learning algorithms and platforms for the recommendation system, ads ranking and search ranking in our company. We also drive substantial impact on core businesses of the company. Responsibilities: 1. Resource Efficiency Optimization in Distributed Orchestration and Scheduling: - Develop and extend distributed orchestration frameworks within the Kubernetes/Godel ecosystem. Select appropriate frameworks based on different business scenarios, and optimize cluster utilization and load balancing strategies according to the specific characteristics of each scenario; - Integrate and expand AutoScaling and automatic parallelization capabilities for various models and tasks. Employ load modeling and analytic methods for different models to automatically optimize resource requests, achieving large-scale improvements in resource usage efficiency and global optimality; - Responsible for preemption and re-scheduling mechanisms for services with different prioritties, and manage automatic resource multiplexing across different clusters and resource types; handle scheduling and load adaptation across multi-datacenter, multi-region, and multi-cloud environments. 2. Building Training System Architecture for Next-Generation Ultra-Large and Ultra-Deep Recommendation Models: - Develop a flexible, elastic and robust distributed training runtime focused on hyper-scaled embeddings and large-scale GPU training; - Design and optimize distributed computing APIs and runtimes geared towards future recommendation and ads model paradigms (e.g., reinforcement learning, fine-tuning and/or distillation); - Collaborate with platform teams to enhance the diagnosability and usability of distributed training systems. 3. Constructing Online Orchestration Architecture for Next-Generation Recommendation Systems: - Build a robust distributed model inference architecture for online learning scenarios involving hyper-scaled embeddings; - Optimize the usability of online recommendation and ads model architectures and MLops workflows.

任职要求

Minimum Qualifications - Bachelor's degree or above, majoring in Computer Science, Engineering or related fields. - Strong programming and coding experience with at least one modern language such as Golang, Python. - Experience contributing to the large scale distributed systems, multi-tenant systems (architecture, reliability and scaling). - Strong analytical abilities and problem solving. - Good communication, self-motivation, engineering practice, documentation, etc. - At least 3 years of relevant experience. Preferred Qualifications - Familiar with large-scale distributed scheduling systems like Kubernetes, Yarn, Flink and/or Spark - Familiar with opensourced orchestration frameworks like VeRL, vLLM, Ray or TFX, etc.

官方来源与核验

字节跳动官方招聘 · 职位编号 7542714368766396680

最近核验:2026-09-16T12:31:41.068845+00:00

查看官方职位详情 ↗

内推申请说明

本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。