← 发现更多职位
字节跳动
实习

Student Researcher (LLM Post Training – Agent & Reinforcement Learning) - 2026 Start (PhD)

面议
美国 · San Jose(圣何塞) · 经验要求见详情
研发实习A253899美国Seed Foundation Model Campus Recruitment - Intern国际招聘

关于这个机会

About the team The Seed LLM Post Training team is responsible for researching cutting-edge posttrain technologies and providing core posttrain capabilities for unified multimodal large models. The team's goal is to research and explore next-generation advanced technologies such as SFT, RM, RL, and self-learning during the posttrain phase, while significantly optimizing and improving key areas including reasoning, coding, agent, and omni model. Responsibilities - Explore large-scale models and optimize systems. - Data construction, instruction tuning, preference alignment, and model optimization. - Improving relevant model capabilities, such as reasoning, code, math etc. - In-depth research and exploration of future use cases.

任职要求

Minimum Qualifications: - Currently pursuing a PhD in Computer Science, AI, or a related field. - Research experience in reinforcement learning, sequential decision-making, or agent behavior. - First-author publications in accredited ML/AI conferences (e.g., NeurIPS, ICLR, ICML). - Solid programming and experimentation skills, including with RL or LLM frameworks. Preferred Qualifications: - Experience with LLM agents, tool use, or prompt-based control. - Familiarity with environments such as WebArena, ALFWorld, or programmatic reasoning tasks. - Understanding of RL techniques such as reward shaping, memory augmentation, or curriculum learning. As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship or any immigration-related benefits.

官方来源与核验

字节跳动官方招聘 · 职位编号 7670325736893565189

最近核验:2026-09-16T14:05:07.210038+00:00

查看官方职位详情 ↗

内推申请说明

本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。