SRE and Observability Engineer
FindErnest · Kolkata, West Bengal, India
IT Services and IT Consulting · 11-50 employees
About the role
Architect and implement unified observability dashboards by integrating data from various monitoring tools and cloud platforms. Manage enterprise infrastructure discovery and automate routine operational tasks using Ansible.
What they look for
Requirements
Requires expert-level knowledge in configuring monitoring tools like Dynatrace, AppDynamics, and Datadog, along with experience in building unified dashboards. Candidates should possess strong process knowledge in ITIL V4 and experience with enterprise tools assessment.
Full description
Key Responsibilities
- Unified Dash boarding: Architect and implement single-pane-of-glass reporting using Grafana/Datadog, integrating data from Splunk, AWS CloudWatch, Azure Monitor, ThousandEyes, Imperva, and Ansible.
- Tools Assessment & Consolidation: Evaluate the current IT landscape, including Datadog, Splunk, Dynatrace, New Relic, AppDynamics, Zabbix, ServiceNow, SMART, SolarWinds, Prometheus, Glowroot, CAMS, etc., to drive consolidation and optimize the target state.
- Enterprise Discovery: Deploy and maintain Device42 for automated infrastructure discovery, application dependency mapping, and continuous CI synchronization with ServiceNow.
- Backend Automation: Implement ticket-based and non-ticket-based routine operational automations leveraging Ansible.
Requirements
Must have Skills:
- Expert-level
knowledge in Monitring Rule setting up and configuring of Dynatrace, App Dynamics, New Relic, Zabbix, Data dogSolarWinds, Prometheus.
- Strong
experience building Unified Dashboards using Grafana ,pulling telemetry from diverse sources including Log/SIEM (Splunk), and Infrastructure.
- Deep experience conducting Enterprise Tools
Assessments and integrating alerts/metrics from legacy APM, network, and infrastructure monitoring tools
- Enterprise Monitoring
& Observability Tools for Onprem & Cloud : (Datadog, New Relic, AppDynamics, Zabbix, SolarWinds Grafana), Log/SIEM (Splunk),
Good to Have Skills:
- Broad
familiarity supporting and assessing monitoring and observability landscapes (e.g., Device42, SMART, Glowroot, CAMS, Thousand Eyes, Thanos, OpsRamp, AWS CloudWatch, Azure Monitor).
- Hands-on
experience with backend automation and orchestration using Ansible, and configuring Major Incident Management (MIM) notifications via PagerDuty.
- ITIL V4
Foundation certification with good process knowledge in Incident, Problem, and Change Management.
- Flexible to work in Off hour / Weekend shift.
Excellent verbal and written communication skills.
Benefits
Locations:Bengaluru, Kolkata/Chennai