← 发现更多职位
字节跳动
Regular

AI Agent Engineer (Stability) - TikTok

面议
新加坡 · Singapore · 经验要求见详情
R&DRegularA60587A新加坡国际招聘TikTok

关于这个机会

Our team is the "TikTok R&D – Service Architecture – Change & Risk Team," responsible for the entire closed loop of TikTok stability spanning "change prevention & control + observability + fault localization/mitigation," and for building Stability AgenticOps as the core productivity infrastructure for the stability domain over the next 1–3 years. On one hand, through change standardization, SLOT canary releases, full-link attribution, and quality inspection, we continuously reduce change-related incidents; at the same time, we build managed batch governance and managed daily releases to drive changes toward unattended operation. On the other hand, focusing on observability high availability, incident recall, and daily alert diagnosis, we ensure core metrics remain observable even under data center failures, surface issues such as effectiveness-related problems and long-cycle low-loss problems as early as possible, and advance diagnosis from "delivering a conclusion" to a "detect–localize–resolve" closed loop. For the AI R&D paradigm, we are not building a "chatty ops assistant"; instead, we build a stability operating system along one horizontal and one vertical axis: horizontally, we accumulate unified context, Skills, planning & orchestration, controlled execution, approval closed loops, and evaluation-driven evolution; vertically, we close the loop in real-world scenarios such as change risk/change hosting, intelligent diagnosis, observability, and incident response. Externally, we serve SREs and business R&D through two forms—CLI and Agent—while high-risk actions retain manual approval. Responsibilities: 1. Responsible for observability, high availability and incident recall; build multi-region disaster recovery capabilities for core observability pipelines such as AppLog/Monitorlog; improve multi-channel detection capabilities across server-side, client-side, and user feedback; and build observability Agents to support scenarios such as core-business impact assessment and SLI lifecycle management. 2. Responsible for the stability of daily alert pipelines, as well as Agent-based intelligent diagnosis capabilities; centered on alert detection and root-cause localization, continuously improve localization accuracy and efficiency. 3. Participate in the development of change-hosting products, including the technical architecture design of internal sub-domains, supporting high-quality and efficient collaboration across sub-domains; through the large-scale batch-change foundation, managed daily business releases, impact analysis and quality inspection (including large models), and the change-hosting Agent (where the Agent understands intent, composes Skills, and maintains collaboration context, while existing release and quality-inspection systems execute and audit), flexibly support governance-type projects and managed daily release scenarios. 4. Participate in the horizontal capability building of Stability AgenticOps, distilling stability platforms, tools, and standard actions into reusable Skills, and building unified context, task planning, controlled execution (dry-run/permissions/approvals), full-link Trace, and evaluation-driven evolution capabilities.

任职要求

Minimum Qualifications: 1. Bachelor's degree or above in Computer Science, Computer Engineering, or other relevant majors. 2. Solid fundamentals and a deep understanding of computer principles; familiarity with operating systems, networking, storage, and related areas. 3. Proficiency in at least one back-end programming language such as C/C++/Go/Python/Java/PHP. 4.. Systematic problem-solving ability, strong communication skills, and a strong sense of ownership; capable of teamwork and communication. Preferred Qualifications: 1. Experience with large-scale distributed systems, stability/SRE, observability, release systems, or change platforms is a plus. 2. Experience with LLM/Agent engineering practice—familiarity with Tool/Skill invocation, context assembly, evaluation, or human-in-the-loop collaboration is a plus.

官方来源与核验

字节跳动官方招聘 · 职位编号 7676368245096368437

最近核验:2026-09-16T12:31:41.068845+00:00

查看官方职位详情 ↗

内推申请说明

本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。