Join us as we inspire
creativity and bring joy to
millions of users worldwide.
@2026 TikTok
Responsibilities
Team Introduction: OSE builds software that helps TikTok engineering teams run services with autonomy and confidence. We are moving from reactive operational support toward self-service reliability platforms, trusted service context, and workflows that scale across the organization. Our work spans the reliability lifecycle: helping teams define quality alerts, route and investigate operational signals, manage incidents, and turn learning into durable improvements. We build platforms that make reliable operations easier for both engineers and on-call teams. OSE sits at the intersection of software engineering and site reliability. We build platform products for monitoring, incident operations, and service metadata. We partner deeply with high-impact engineering teams and design automation that remains accountable to the people who operate production systems. Engineers also embed with partner teams through a scheduled 12-hour on-call rotation. This builds direct understanding of live operational workloads and informs better platform design. The rotation is no more frequent than once every four weeks. We are building toward AI reliability workflows. In this role, you will help make engineering and operational work more agent-ready through clear product contracts, safe automation, and evidence-driven iteration. AI can assist with triage, drafting, and bounded actions. Engineers remain responsible for validation, production decisions, and outcomes. - Build and evolve reliability platforms that engineers use daily: APIs, workflows, configuration surfaces, and operator experiences. - Improve the alarm lifecycle from authoring through triage, handling, remediation, and feedback into better rules and metadata. - Design for safe automation by applying appropriate dry-run paths, approval gates, rollback options, and traceability to agent-assisted or scripted actions. - Keep humans accountable for policy decisions, high-risk changes, and production outcomes. Agents can propose and assist. People own merges, rollouts, and incident resolution. - Validate before you ship: tests, static checks, staged rollout, and self-verification evidence where the platform supports it. - Instrument workflows so adoption, quality, and failure modes are observable without confusing activity with causality. - Partner with internal teams to prove patterns, gather feedback, and hand off repeatable workflows so partners can self-serve. - Embed with partner teams through a 12-hour on-call rotation to understand live operational workloads and translate recurring friction into platform improvements. Rotations occur no more frequently than once every four weeks. - Participate in reliability culture: incident learning, post-incident follow-up, on-call hygiene, and documentation that downstream teams and agents can trust. - Contribute to AI platform capabilities where appropriate: agent-usable skills, product contracts, read-only query paths, and governed write paths with clear permissions.
Qualifications
Minimum Qualifications: - Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. - Professional experience building or operating backend services, platform tooling, or infrastructure-facing products. - Proficiency in at least one backend language (for example Go, Java, or Python) and comfort reading code across a polyglot stack. - Understanding of distributed systems basics: services, APIs, deployments, observability, and failure modes & Familiarity with on-call concepts: alerting, incident response, escalation, and runbooks. - Ability to reason about data models, ownership, and correctness when metadata drives routing or automation. - Write clear, reviewable code with tests appropriate to risk & Use version control, code review, and CI/CD as normal practice & Debug production issues with logs, metrics, traces, and structured hypotheses. Preferred Qualification: - More than 1 years relevant work experience from a large-scale internet business -Stay curious about AI-assisted development without overstating what models can safely own in production.
Job Information
About TikTok
TikTok is the leading destination for short-form mobile video. At TikTok, our mission is to inspire creativity and bring joy. TikTok's global headquarters are in Los Angeles and Singapore, and we also have offices in New York City, London, Dublin, Paris, Berlin, Dubai, Jakarta, Seoul, and Tokyo.
Why Join Us
Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day.
We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.
Diversity & Inclusion
TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.