← 发现更多职位
字节跳动
正式

Site Reliability Engineer - Traffic Infrastructure

面议
新加坡 · Singapore(新加坡) · 经验要求见详情
研发基础架构正式A154583A新加坡国际招聘

关于这个机会

About the Team The Traffic Infrastructure team leverages unified platform capabilities to manage global edge infrastructure (China & Non-China), both self-built and third-party, providing standardized, compliant, scalable, and cost-effective traffic infrastructure capabilities for edge services. Our vision is to build a global edge traffic infrastructure platform and become the long-term cornerstone of ByteDance’s global edge business in terms of scale, performance, and cost. Responsibilities - Responsible for the operation and maintenance as well as stability assurance of ByteDance's "Network-Traffic Infrastructure". - Responsible for the delivery and operation & maintenance of the production system, including the delivery, change and release of facilities, components and products, and improving the efficiency of both delivery and operation & maintenance. - Responsible for the design and implementation of the stability assurance system, covering system observability (monitoring/alerts/logging), troubleshooting (root cause analysis/impact assessment), and issue resolution (manual/self-healing). - Responsible for the design and implementation of the emergency response system, including work such as ticket processing, emergency response, risk governance, and long-term optimization, to enhance the risk emergency response capability.

任职要求

Minimum Qualification(s) - Bachelor's degree or above in computer science or a related field, with at least 3 years of relevant experience in R&D, system operation and maintenance, or SRE. - Familiar with infrastructure architecture, and have a solid understanding of Kubernetes, edge computing, cloud networking, Load Balance, microservice architecture and other related technologies. - Possess strong analytical skills, excellent communication abilities, a strong sense of responsibility and team spirit. Preferred Qualification(s) - Possess practical operation and maintenance and stability assurance experience in Kubernetes, cloud computing, edge computing, and cloud networking. - Well-versed in high availability, stability assurance, and emergency response systems for infrastructure or distributed systems, with relevant operation and maintenance experience. - Hands-on experience in handling risks, potential hazards and failures of infrastructure or distributed systems.

官方来源与核验

字节跳动官方招聘 · 职位编号 7665587592421820677

最近核验:2026-09-16T14:05:07.210038+00:00

查看官方职位详情 ↗

内推申请说明

本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。