← 发现更多职位
字节跳动
Regular

Senior Research Engineer - AI-Native Online Datastore Systems

面议
美国 · San Jose · 经验要求见详情
R&DRegularA192131美国国际招聘TikTok

关于这个机会

We are TikTok's Online Datastore Systems team, responsible for designing, building, and governing the core online storage infrastructure that powers our global business—including databases, caches, data synchronization, metadata management, storage governance, and global data distribution. Our mission is to deliver online data storage services with ultimate performance, reliability, and intelligence for hundreds of millions of users and countless business scenarios worldwide. The team is undergoing a pivotal paradigm shift: from manual oncall and hand-crafted governance toward AI-native development and operations. We already have early practices in production—an Oncall Agent (knowledge base + runbooks + tiered write-permission model), a storage replica governance Skill, and an Agent Box toolchain—but we need someone to design the AI infrastructure layer from scratch, establish an evaluation framework, and drive org-wide adoption. This is a greenfield opportunity. You won't be maintaining existing AI systems—you'll be the team's first dedicated AI-Native Research Engineer, balancing research exploration with hands-on engineering: studying how agents work best in online data storage contexts, and turning those insights into reusable infrastructure and workflows. Here, you will tackle world-class challenges in globalization, multi-region active-active, compliance, and cost efficiency—while using AI to redefine how storage systems are built and operated. Responsibilities - Design and build the AI Context & Knowledge Layer: Architect a centralized context layer for the online datastore systems team, integrating knowledge bases, runbooks, code repositories, and real-time system state so AI agents have grounded, traceable, team-specific domain knowledge. Continuously iterate on knowledge structure and retrieval strategies to improve answer quality and evidence-chain completeness. - Research and implement Agentic Workflows: Explore agent applications across the development lifecycle (AI-assisted coding, automated testing, PR pre-review, deployment validation) and the operations lifecycle (oncall triage, structured troubleshooting, storage governance task submission). Design multi-agent orchestration frameworks with tiered permission and safety guardrails (read-only / confirm-then-execute / human-execute). - Establish Agent Evaluation & Experimentation: Design offline eval sets and core metrics (hit rate, evidence-chain completeness, misoperation rate). Validate agent effectiveness through shadow mode or A/B experiments. Build a reproducible evaluation methodology to drive continuous improvement. Share findings through tech talks or internal write-ups. - Storage Platform Skill-ification & Toolchain Integration: Encapsulate core online storage capabilities (metadata queries, workflow troubleshooting, storage replica governance, DDL changes, etc.) as reusable Skills. Integrate with gdpa-cli, lark-cli, and real-time query tools. Build standardized AI development environments. - Drive AI-Native Adoption & Enablement: Accelerate team-wide adoption through pairing, workflow demos, architecture reviews, and best-practice documentation. Balance AI-driven speed with code quality and system safety. Serve as the technical advocate for AI-native transformation.

任职要求

Minimum Qualification(s): - Research background (required — one of the following): - PhD in Computer Science, Artificial Intelligence, or a related field; or - Thesis-based Master's degree with a research focus in AI, ML, software engineering, or systems. - Proficiency in Python and Go (or C++ / Java), with the ability to translate research prototypes into production systems; AI-Native coding mindset. - Deep expertise in LLM application engineering: RAG, tool use / function calling, structured outputs, prompt / context engineering; proven track record of engineering research into shipped systems. - Solid distributed systems foundation; familiarity with online storage, caching, data sync, or metadata management is a plus. Preferred Qualification(s): - Peer-reviewed publications or top-tier conference acceptances in agent systems, LLM applications, AIOps, or SE4ML. - Research-grade experience designing agent evaluation, offline benchmarks, or experimentation frameworks. - Multi-agent orchestration and guardrail design for high-risk operations. - Global online storage infrastructure background (multi-region active-active, cross-region sync, compliance governance).

官方来源与核验

字节跳动官方招聘 · 职位编号 7658238613570718005

最近核验:2026-09-16T12:31:41.068845+00:00

查看官方职位详情 ↗

内推申请说明

本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。