Senior Site Reliability Engineer at Next Ventures
Dubai, Dubai, United Arab Emirates -
Full Time


Start Date

Immediate

Expiry Date

15 Dec, 26

Salary

0.0

Posted On

16 Sep, 26

Experience

0 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description

What You Bring

  • 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or strongly reliability-focused DevOps work.
  • Strong hands-on experience with centralized log management platforms — ingestion, parsing, structured logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
  • Able to diagnose incidents through log analysis across both Linux and Windows environments, isolating root causes under pressure.
  • Designs alerting systems with tiered thresholds that minimize noise while catching real problems early, with clear escalation paths.
  • Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
  • Experienced defining and operating SLOs, SLIs, and error budgets for customer-facing services.
  • Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
  • Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
  • Hands-on experience operating services on Kubernetes, EKS preferred.
  • Familiar with Infrastructure-as-Code tooling such as Terraform for collaboration with the DevSecOps team.
  • Evidence-driven — profiles and measures before guessing; validates every optimization against before/after data.
  • Reliability-oriented — treats detection speed and signal quality as first-class engineering problems.
  • Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.

How To Apply:

Incase you would like to apply to this job directly from the source, please click here

Responsibilities
Loading...