Aug 7, 2026
Electric Energy Jobs

Senior Systems Reliability Engineer

Organization:
Electric Reliability Council of Texas
Region:
Canada, Texas, Taylor
End of contest:
November 4, 2026
Type:
Full time
Category:
Systems engineer
Description
JOB SUMMARY

The Senior Systems Reliability Engineer applies software engineering discipline to reliability problems designing, building, and operating the systems that make production software measurable, scalable, and self-healing. This role treats operational challenges as engineering problems: when a process is manual, it gets automated; when a failure mode is unknown, it gets instrumented; when a system degrades, the degradation is understood before it recurs.

At this level, the specialist owns SLO and error budget frameworks for assigned systems, architects the observability stack that the team relies on, leads engineering-driven incident response, and holds NERC/CIP compliance responsibility for assigned systems. This role partners directly with Software Engineers as a technical peer participating in design reviews, influencing architecture decisions for reliability, and building the production readiness standards that govern how software ships. Advancement to Lead is based on demonstrated ability to define reliability engineering standards at the platform level, influencing practice across multiple teams and portfolios.

JOB DUTIES

  • Performs complex reliability engineering work autonomously; recognized subject matter expert within the team and adjacent teams.
  • Designs and builds production software systems, reliability tooling, and automation frameworks; treats operational problems as engineering problems to be solved through code.
  • Owns SLO governance, error budget management, and observability architecture for assigned systems; leads engineering-driven incident response including failover scenarios.
  • Holds NERC/CIP compliance responsibility for assigned systems; formally mentors less experienced specialists; may coordinate team delivery and on-call activities.

ADDITIONAL JOB DUTIES

Core Expectations

The following expectations apply at all Systems Reliability Specialist levels. Scope and independence expand with each level.

  • Engineer reliability solutions: when a process is manual and repeatable, automate it; when a failure mode is opaque, instrument it; when a system is fragile, redesign the failure boundary.
  • Define and own SLIs and SLOs for assigned systems; treat error budgets as a shared engineering contract with development teams, not an operations metric.
  • Respond to production incidents as an engineer: form a hypothesis, isolate the failure, resolve it, and close the loop with a post-mortem that addresses root cause.
  • Instrument systems so that on-call responders have sufficient telemetry to diagnose and act without tribal knowledge.
  • Participate in 24/7 on-call rotation; treat every alert as signal either actionable or worth eliminating.
  • Write production-quality code: reliability tooling, automation frameworks, and operational software are held to the same engineering standards as application code.
  • Partner with development teams as a peer in design reviews; reliability is designed in, not bolted on after deployment.

Reliability Engineering

Senior specialists design and build the engineering systems that make production software reliable. This is software engineering applied to operational problems the output is code, frameworks, and automated systems, not tickets and runbooks alone.

  • Design, build, and maintain reliability tooling: automated remediation systems, self-healing infrastructure components, and operational software that reduces human intervention in production.
  • Own SLO and error budget definitions for assigned systems; review error budget consumption with development teams and drive engineering decisions based on budget status.
  • Architect and implement chaos engineering programs: define failure injection scenarios, automate resilience tests, and validate recovery behavior against defined SLOs.
  • Build and maintain CI/CD reliability gates: automated canary analysis, progressive delivery validation, and rollback triggers based on SLI thresholds.
  • Design capacity planning models for assigned systems; build tooling to project resource needs and surface capacity risks before they affect availability.
  • Contribute to production readiness reviews: define and enforce the engineering criteria that a system must meet before it ships to production.
  • Reduce operational toil through engineering: measure toil, track reduction targets, and build the automation that eliminates it.

Observability & Instrumentation

Observability is an engineering discipline. Senior specialists design and build the telemetry systems that make production behavior understandable not just monitored.

  • Architect MLTP (Metrics, Logs, Traces, Profiling) observability solutions using the Grafana LGTM stack (Loki, Grafana, Tempo, Mimir), Dynatrace APM, Splunk, and Datadog.
  • Define and enforce instrumentation standards: structured logging schemas, metric naming conventions, trace context propagation, and continuous profiling configuration for assigned systems.
  • Build distributed tracing coverage across service boundaries; identify and close observability gaps that produce blind spots during incidents.
  • Design SLI instrumentation: translate user-facing reliability requirements into specific, measurable signals that accurately represent system health from the user's perspective.
  • Build and maintain alerting frameworks: alerts must be actionable, calibrated to SLO burn rate, and free of noise; own alert quality as an engineering output.
  • Correlate application performance data JVM heap behavior, GC pressure, thread contention with infrastructure events to enable root cause analysis across layers.

Incident Response & Problem Management

Incident response at this level is an engineering activity. Senior specialists lead the technical response to high-severity events, own the post-mortem process, and drive the engineering work that prevents recurrence.

  • Lead high-severity incident response for assigned systems, including dual-datacenter failover execution; own the technical resolution from detection through remediation.
  • Apply structured root cause analysis: distinguish symptoms from causes, identify contributing factors across system layers, and drive remediation that addresses root cause rather than surface behavior.
  • Author post-mortems that produce actionable engineering work items not process improvements alone; track remediation to completion and validate effectiveness.
  • Diagnose complex cross-layer failures: Java/JVM application failures, distributed system race conditions, database connection pool exhaustion, messaging system backpressure, and cross-datacenter synchronization issues.
  • Build and maintain incident response runbooks as engineering artifacts: automated where feasible, version-controlled, and validated during chaos engineering exercises.
  • Participate in blameless post-mortem facilitation; model the engineering culture that treats incidents as system failures, not human failures.

Read the full posting.

Contact

Electric Reliability Council of Texas





Texas United States

www.ercot.com