operational region: birmingham · gb discipline: sre ← platformsignals.dev
--:--:-- BST
profile / sowmya-shree · v15.0

Sowmya
Shree

Site reliability engineer. Fifteen years inside Tier 1 investment-banking infrastructure — measured the way an SRE measures anything: with SLIs, SLOs, and an error budget.

Status
Open
Choosing next team
Uptime
15+ years
Continuous service
Tier
Tier 1 · IB
Investment banking
On-call
Veteran
Multi-year primary
● Next move
Current status: Deliberately choosing the next team. Open to SRE and Reliability Engineering roles where reliability is treated as a feature — the right problem, the right culture, UK-based, hybrid or remote.

The short version.

Fifteen-plus years inside one of the largest investment banks in the world. The work has consistently been the unglamorous and important kind — keeping production systems honest, building the observability other engineers depended on, and making the platforms that let teams ship safely.

I came into SRE the long way. Production support taught me how systems actually fail at three in the morning. Platform work taught me how to stop most of those failures from happening in the first place. SRE gave me the discipline and the language to put numbers on both. Along the way I've designed and operated end-to-end services — including Data Lake and Lakehouse platforms — and navigated the compliance and reporting requirements that come with working in regulated financial infrastructure. The page below describes how I operate the way I'd describe any other production service.

How I measure myself.

The indicators I track on my own work, with the targets I hold myself to. These aren't aspirational — they're the operating standards I've maintained across multiple production teams, and they're what I'd bring to a new one.

IDIndicatorTargetStatus
SLI-01 Incident responseAcknowledge within rotation SLA, lead remediation, restore service before chasing root cause ≥ 99% Meeting
SLI-02 Postmortem turnaroundBlameless postmortem published with actions assigned and owned ≤ 5 days Meeting
SLI-03 Toil reductionTime spent on repetitive manual work, measured per quarter ≤ 50% Meeting
SLI-04 Observability coverageServices I own have SLIs, dashboards, and alerts routed to a real on-call 100% Meeting
SLI-05 Knowledge transferRunbooks current, on-call rotation cross-trained, tribal knowledge written down No single point Meeting
SLI-06 Continuous learningOne substantive piece of public writing or internal deep-dive per quarter ≥ 4 / yr In budget

Reliability against target.

Rolling 12-month attainment against the SLIs above. Amber line marks target.

Incident response
SLI-01 · target 99%
99.4%
Postmortem turnaround
SLI-02 · target ≤ 5d
3.2d avg
Toil reduction
SLI-03 · target ≤ 50%
~35%
Observability coverage
SLI-04 · target 100%
100%
Continuous learning
SLI-06 · target 4/yr
3 / 4 YTD

Figures are honest self-reported approximations across recent production teams, not audited metrics — same caveat as any internal SLO report.

What I won't trade.

An error budget defines what a service is allowed to spend reliability on. Mine works the same way — these are the trade-offs I won't make, because making them puts the work itself at risk.

EBP-01No silent production changes.Every change in prod has a record, a ticket, and a backout. Saving thirty minutes today by skipping change control is a debt I'll be paying back at 2am next week.
EBP-02No blameful postmortems.A postmortem that names a person rather than a failure mode teaches the team to hide problems. I'll push back on this every time, including with leadership.
EBP-03No undocumented heroics.If only one person can fix the system, the system is broken even when it's working. Knowledge that lives in one head is a SEV waiting for a holiday.
EBP-04No "monitor it later."If a service ships without SLIs and a real on-call destination, it isn't in production — it's in someone's career risk profile.
EBP-05No reliability theatre.Dashboards nobody reads, alerts nobody actions, runbooks nobody updates — these are worse than nothing because they create the illusion of safety. I'd rather have less, and have it work.

Production services I've owned.

Three domains inside Tier 1 investment-banking infrastructure. Different scale bands, different failure modes, same craft.

svc / trading-platform● tier 0 · mission critical

Trading & Front-Office Systems

Reliability engineering for systems where seconds have a P&L attached to them. The masterclass in why latency tails matter, why "mostly working" isn't a state, and why your monitoring needs to be faster than your traders.

Scale
High-throughput
Criticality
Revenue-impacting
On-call
24×7 primary
svc / observability-platform● tier 1 · infrastructure

Observability & Monitoring Tooling

Building and operating the internal tooling other engineers depended on to know what their services were doing. Metrics pipelines, dashboards, alerting infrastructure — the systems that make other systems debuggable.

Scale
Bank-wide
Users
Engineering org
Failure mode
Blast-radius critical
svc / internal-developer-platform● tier 1 · platform

Internal Developer Platform

Self-service infrastructure for engineering teams who shouldn't have to know what a VPC is to ship a service. Paved roads, golden paths, sensible defaults — making the right thing the easy thing.

Scale
Multi-team
IaC
Terraform · AWS
Customer
Internal engineering

The stack I work in.

Tools I've used in anger, in production, at scale. Not a wishlist.

Language & IaC
PythonJavaTerraformBashLinuxYAML (sadly)
Cloud
AWS
Observability
DatadogPrometheusSplunkOpenTelemetry
Data Platform
Data LakeLakehouseEnd-to-end service design
Operations & Compliance
ServiceNowCompliance & reporting24×7 on-callPostmortem facilitationIncident command

Reach me.

If you're hiring for SRE or Reliability Engineering — or you just want to compare notes on the craft — LinkedIn is the fastest path.

BaseBirmingham, UK · hybrid / remote
Availability● Choosing next move — responding within 24h