logologo
Hunt UK Visa Sponsors
Jobs
logologoHunt UK Visa Sponsors

Find jobs from UK licensed visa sponsors — Companies House verified, updated daily.

About

How does it workContact Us

Find Work

JobsJobs by RoleLicensed SponsorsVisa TypesSponsor Statistics

Resources

BlogGlossaryOccupation EligibilityIncome Tax CalculatorILR Tracker

Content on this site is for general information only and does not constitute legal advice. Always consult a regulated UK immigration solicitor for advice specific to your situation.

Copyright © 2026. All rights reserved.

Entrust

Principal Software Reliability Engineer, AI/ML Product Engineering

CompanyEntrust
LocationLondon Area, United Kingdom
Posted At3/3/2026

UK Visa Sponsorship Analytics

Analytics are greyed out due to low classification confidence (39.0%).
Occupation Type
Programmers and software development professionals
Occupation Code Skill LevelHigher Skilled
Sponsorship Salary Threshold
£54,700 (£28.05 per hour)
Occupation rate applies

Above analytics are generated algorithmically based on job titles and may not always be the same as the company's job classification. You can also check detailed occupation eligibility, and salary criteria on our UK Visa Eligible Occupations & Salary Thresholds page.

Disclaimer: Hunt UK Visa Sponsors aggregates job listings from publicly available sources, such as search engines, to assist with your job hunting. We do not claim affiliation with Entrust. For the most up-to-date job details, please visit the official website by clicking "Apply Now."

Description

About This Role

This is a Product Reliability position, not an infrastructure SRE role. Our DevOps team manages the infrastructure platform; this role focuses on application and service-level reliability, working directly with product engineers.

This is the first role of its kind in product engineering. Reporting to the VP of Product Engineering for Consumer Identity, you’ll drive reliability efforts across the team: defining the roadmap, prioritizing initiatives, and partnering with engineering directors and senior ICs to deliver them.


Why Join Us

  • Greenfield opportunity: You’ll define Product Reliability as a discipline here. Build the playbook, not inherit one.
  • High-impact domain: Consumer Identity powers identity verification and biometric authentication for some of the world’s largest financial institutions. Our reliability directly impacts fraud prevention and customer onboarding at scale.
  • Real authority: Direct line to VP Engineering, budget for tooling, seat at architecture council and service reviews.
  • Strong foundation: We’re not firefighting. 99.98% uptime means you’re optimizing, not triaging chaos.
  • Technical depth: Work across ML pipelines, computer vision systems, and mobile SDKs (not just YAML and dashboards).
  • Ownership culture: Engineers own their services end-to-end; you’ll amplify that, not replace it

  • Experience Level

    Staff SRE

    • 8+ years in software engineering
    • 4+ years in reliability/SRE
    • Drives reliability initiatives across multiple teams; hands-on with complex systems

    Principal SRE

    • 15+ years in software engineering
    • 6+ years in reliability/SRE
    • Sets technical direction org-wide; influences business-unit-level reliability strategy

    We’re open to either level. Scope and compensation will match your experience. Principal candidates should demonstrate cross-org impact and a track record of building reliability programs from scratch.


    Current State Incident Analysis (2020–2025)

    • Postmortem volume peaked in 2023, down 48% since then despite increased release cadence
    • P0-to-P1+ ratio remains stable despite lower overall incident volume
    • 65% change-induced incidents (deployments, migrations, config changes); 35% organic (third-party outages, expirations, attacks)
  • Change-induced ratio improved modestly: 69% → 62%
  • Detection time: 35 min → 18 min
  • Customer-first detection: 40% → 22%

  • Availability Targets

    • 2024 & 2025 average uptime: 99.98% (as available in our public status page)
    • Goal: Consistent 99.99% (four nines) average uptime, SLO breach reductions


    System Simplification

    • We’re reducing system complexity to narrow the reliability target area:
    • Microservices (K8s deployments/rollouts) reduced 29% from peak, with further cuts planned for 2026
    • Goal: Smaller footprint, higher reliability, lower cost for new regions


    Role Objectives

    • Primary goal: Improve release safety, reduce releases that cause downtime or SLO degradation.
    • We already have foundational systems in place:
    • Automated test coverage and crowd testing
    • A/B testing and dark canaries
    • Progressive rollouts (infrastructure and application level)
    • Back-testing against historical data
    • To consistently exceed four nines, we need to mature these systems and build new capabilities.


    Ideal Candidate Profile

    Mindset

    • Passionate about reliability as a discipline, not just a checkbox
    • Focused on reliability, not product features, but willing to learn the product to understand impact
    • Hands-on: eager to build tooling and systems
    • Pragmatic about balancing reliability with development velocity


    Required Skills

    Software Engineering

    • Strong software engineering in at least one of our backend languages (Python, Ruby, Node.js); able to navigate most of our codebase
    • Experience building reliability tooling: progressive delivery, automated rollbacks, monitoring/alerting

    Reliability Patterns

    • Deep knowledge of resilience patterns: circuit breakers, bulkheads, back-pressure, retries with backoff, rate limiting, load shedding, graceful degradation
    • Solid incident management and blameless postmortem practices

    Observability

    • Proficiency with observability: distributed tracing, structured logging, metrics instrumentation
    • Uses data to drive decisions: experienced with SLIs, SLOs, and error budgets

    Communication

    • Skilled at influencing without authority
    • Able to hold deep technical reliability discussions with senior ICs


    Nice-to-Have

    • Experience with chaos engineering (fault injection, game days, controlled failure experiments)
    • ML system reliability experience (mixed I/O and CPU-bound workloads, non-deterministic behavior, model serving)
    • Familiarity with our specific stack (Datadog, Kubernetes, AWS, GitLab CI/CD)
    • Experience leveraging LLMs for code analysis, design doc review, or automated runbook generation
    • On-call experience in a high-availability environment


    Our Stack

    • Backend: Python, Ruby on Rails, Node.js
    • Frontend: React, TypeScript
    • Mobile: Swift (iOS), Kotlin (Android), React Native
    • Infrastructure: AWS, Kubernetes, Terraform, SNS, SQS
    • Databases: PostgreSQL (Aurora), Redis, OpenSearch
    • Observability: Datadog, Splunk, Sentry
    • ML: PyTorch, TensorFlow
    • CI/CD: GitLab (on-prem)