Site Reliability Engineer, Recommendation Architecture (ByteDance Singapore)
关于这个机会
About The Team Our Recommendation Architecture Team is responsible for building up and optimizing the architecture for our recommendation system to provide the most stable and best experience for our users. On the SRE team of Recommendation Architecture, you'll have the opportunity to sharpen your expertise in coding, performance analysis, large-scale system operation, and get heavily involved in the process of hardware/capacity decision-making. SRE ensures that the recommendation services at ByteDance have the highest level of availability, as well as creating highly automated systems and pipelines. Responsibilities - Reliability and operation optimization for large-scale clusters of Recommendation System. - Continuous integration and delivery of core services, optimizing the efficiency and automation of operation, and improving service stability and R&D efficiency. - Cloud platformization, resource optimization and SLA guarantee for large-scale clusters. - Collaboration with software engineer to design and implement DevOps solutions to Improve the efficiency of the entire R&D process. - Research, design, and develop computer and network software or specialised utility programs. - Analyse user needs and develop software solutions, applying principles and techniques of computer science, engineering, and mathematical analysis. - Update software, enhances existing software capabilities, and develops and direct software testing and validation procedures. - Work with computer hardware engineers to integrate hardware and software systems and develop specifications and performance requirements.
任职要求
Minimum Qualifications - Bachelor's degree or above in computer science, software engineering, or a related field - Good programming experience with at least one of the following languages: Shell/Python/Perl/Go/C++. Preferred Qualifications - Operation experience of large-scale systems, familiar with system operation skills on Linux and network. - Expertise in analyzing, and troubleshooting large-scale distributed systems.
内推申请说明
本站为独立内推协助平台。申请会交由管理员核实岗位与内推渠道,不等于已在公司官网投递;薪资、岗位状态和实际招聘流程以官方信息为准。