jobs@entomo
Join us to build enterprises of tomorrow
Senior SRE
Location: null
Tech Stack: Java, Spring Boot, Angular/React, Kubernetes
Role Summary
We are seeking a Senior Software Engineer (Reliability & Observability) who is a strong hands- on Java/Spring Boot engineer and takes end-to-end ownership of production reliability.
This role moves beyond traditional monitoring and ticket escalation. The engineer is expected to proactively identify production issues, perform in-depth debugging at the application code level, implement permanent fixes, and safely deploy improvements to production systems.
The position involves approximately 70% backend application development and 30% reliability engineering/SRE and DevOps responsibilities, with a strong emphasis on observability, incident response, and system resilience.
Key Responsibilities
Application Development & Incident Resolution (≈70%)
Own production issues from detection through resolution, ensuring minimal impact to users and systems.
Debug and resolve issues directly within Java and Spring Boot codebases, including but not limited to:
○ Performance bottlenecks and latency issues
○ Memory leaks and JVM tuning problems
○ Concurrency and multithreading issues
○ API failures and error-handling gaps
Database and caching inefficiencies
Conduct deep root cause analysis (RCA) for production incidents and implement long- term, sustainable fixes, rather than temporary workarounds.
Strengthen application resilience by designing and implementing:
○ Robust exception-handling strategies
○ Timeouts, retries, and circuit breaker patterns
○ Graceful degradation and fail-safe mechanisms
Develop clean, maintainable, and production-grade code, supported by strong unit, integration, and reliability-focused testing.
Reliability, Observability & DevOps (≈30%)
Design, implement, and continuously improve monitoring, alerting, and observability solutions across services.
Analyze metrics, logs, and distributed traces to proactively detect and prevent reliability issues.
Apply SRE best practices to improve system stability and availability, including:
○ Defining and tracking SLIs, SLOs, and error budgets
○ Leading or contributing to incident postmortems and driving preventive action plans
Collaborate with CI/CD systems to safely deploy fixes and improvements to production.
Support and optimize cloud-native and containerized workloads using Docker and Kubernetes.
Role Expectations
This is not a monitoring-only or escalation-focused role.
The engineer is expected to fix production issues directly in code and take accountability for system reliability.
Success in this role depends on a strong engineering mindset, ownership, and a proactive approach to production stability.

