← 发现更多职位
字节跳动
正式

Site Reliability Engineer - System Service Global

面议
美国 · San Jose(圣何塞) · 经验要求见详情
研发基础架构正式A231689美国国际招聘

关于这个机会

The Global System Service team owns the infrastructure services and management solutions that power ByteDance's data centers outside of China — from day-to-day operations to long-term architecture design and maintenance. The team specializes in composing end-to-end solutions by drawing on both open-source community tools and in-house developed products, tailored to both the business requirements and the operational complexities of large-scale infrastructure across ByteDance's non-China regions. Our mission is to deliver efficient infrastructure solutions and a stable, secure system environment for ByteDance's global business. We are looking for a self-motivated system engineer that is equipped with SRE mindset and DevOps skills. Your responsibilities will include: - Manage and maintain large-scale host infrastructure across ByteDance's non-China data centers, covering OS lifecycle management, configuration standardization, and fleet-wide health monitoring. - Own the reliability and availability of core data center foundational services, including DNS, NTP, DHCP, NAT, APT repository, and Kerberos authentication. - Design and implement deployment architectures for foundational services, ensuring high availability, fault tolerance, and disaster recovery across regions. - Develop and enforce SLOs for managed services; lead incident response, root cause analysis, and post-mortem reviews to drive continuous reliability improvements. - Collaborate with network, security, and application teams to ensure foundational services meet the evolving demands of global business growth. - Identify automation opportunities across host management and service operations; drive tooling and process improvements to reduce toil and increase operational efficiency.

任职要求

Minimum Qualifications: - Bachelor’s degree or higher in Electrical Engineering, Computer Engineering, Computer Science or related majors. - Solid experience in large-scale Linux host management, including OS deployment, configuration management, patching, and fleet operations. - Strong hands-on knowledge of core data center foundational services: DNS (BIND/PowerDNS), NTP, DHCP, NAT, APT repository management, and Kerberos. - Proficiency with DevOps tooling, including configuration management tools (e.g., Ansible, Salt, Puppet) and CI/CD pipelines. - Familiarity with SRE principles and practices, including SLO/SLI definition, error budget management, and blameless post-mortems. - Solid understanding of high availability design patterns, active-active/active-passive architectures, and disaster recovery strategies. - Strong troubleshooting skills across the Linux system stack and network layer. Preferred Qualifications: - Experience managing host fleets at scale (thousands of nodes or above) in a production environment. - Scripting or development experience in Python, Go, or Bash for automation and tooling. - Exposure to hybrid or multi-region data center environments.

官方来源与核验

字节跳动官方招聘 · 职位编号 7639133243979827509

最近核验:2026-09-16T14:05:07.210038+00:00

查看官方职位详情 ↗

内推申请说明

本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。