Wallstreetdocs Ltd

Site Reliability Engineer

Wallstreetdocs Ltd · Dhaka, Dhaka Division, Bangladesh

Information Technology & Services · 201-500 employees

19 h ago
Mid (2-5 yrs) Full-time Bangladesh
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Proactively identify, diagnose, and resolve application reliability issues across Node.js and Java environments. Collaborate with cross-functional teams to enhance system resilience, lead incident reviews, and maintain comprehensive monitoring solutions.

What they look for

Linux Windows Node.js Java Grafana Elastic APM Zabbix Proxmox Kubernetes AWS Troubleshooting Incident Management System Performance Analysis Root-cause Analysis Infrastructure Monitoring Communication

Requirements

Requires a minimum of 3 years of experience in troubleshooting application-level issues in Linux and Windows environments. Candidates must possess hands-on experience with Node.js, Java, and various monitoring tools like Grafana, Kubernetes, and AWS.

Benefits

Provident Fund Health Insurance Festive Bonuses Flexible Annual Leave Policy Parental Leave Breakfast Lunch Unlimited coffee and snacks Gaming Zone

Full description

We’re seeking a Solution Reliability Engineer who thrives on solving intriguing technical challenges and ensuring the continuous reliability of customer-facing applications. This role is ideal for a technically versatile individual, a true all-rounder comfortable navigating between infrastructure, DevOps, and client facing support ensuring our solutions remain consistently stable, efficient, and available.

As our Solution Reliability Engineer, you won't be limited by specific technologies or platforms; instead, you'll use your analytical skills, curiosity, and technical intuition to dive into diverse environments, understand complex problems, and deliver reliable solutions.

Our infrastructure is predominantly Linux-based with some Windows environments, hosting applications primarily developed in Node.js and Java. We extensively utilise tools such as Grafana, Elastic APM, Zabbix, Proxmox, Kubernetes, and AWS. You’ll collaborate closely with our DevOps, Infrastructure, and Development teams, making a real difference in how we deliver excellent service to our customers.

Key Responsibilities

  • Proactively identify, diagnose, and resolve application reliability issues across various technology stacks, primarily focused on Node.js and Java applications.
  • Collaborate across DevOps, Infrastructure, and Development teams to enhance solution resilience and stability.
  • Provide clear, approachable, and effective communication with stakeholders from technical and non-technical backgrounds.
  • Analyse system performance data, identify trends, and suggest improvements to prevent recurring issues.
  • Lead root-cause analyses, post-mortems, and incident reviews, fostering a culture of continuous learning and improvement.
  • Document solutions, troubleshooting steps, and system configurations to empower cross-team knowledge sharing.
  • Drive enhancements and ensure comprehensive coverage of both conventional infrastructure monitoring and advanced application performance monitoring solutions.
  • Leverage and continuously improve our monitoring stack, including Grafana, Elastic APM, Zabbix, and other tools.

Qualifications and Experience

  • Demonstrated minimum 3 years of experience in troubleshooting and resolving application-level issues across diverse environments (Linux and Windows).
  • Broad technical knowledge with the confidence and curiosity to rapidly learn new systems and technologies.
  • Familiarity and hands-on experience with Node.js and Java-based applications.
  • Experience with monitoring and observability tools such as Grafana, Elastic APM, and Zabbix.
  • Understanding of virtualization and container technologies, particularly Proxmox and Kubernetes.
  • Experience working with AWS environments (optional)
  • Strong analytical, diagnostic, and problem-solving skills with a structured approach to incident management.
  • Excellent interpersonal and communication skills, capable of effectively engaging stakeholders across various teams.
  • Comfortable operating in complex environments involving multiple applications, services, and technologies.

Personal Attributes

  • Curious, sense of ownership, approachable, and proactive in exploring new challenges.
  • Collaborative team-player with a positive and open communication style.
  • Resilient and adaptable, able to confidently handle pressure during incidents and emergencies.
  • Committed to continuous improvement, both personally and organizationally.

Benefits & Perks:

  • Competitive compensation package 
  • Provident Fund 
  • Health Insurance 
  • 2 Festive Bonuses 
  • Yearly Review 
  • Flexible Annual Leave Policy + All BD Govt. approved leaves 
  • Parental Leave (maternity & paternity) 
  • Opportunity to work with our incredible global teams across multiple time zones 
  • Breakfast, lunch, unlimited coffee & snacks in office 
  • Gaming Zone & recreational facilities 
  • Supportive, collaborative, and inclusive work environment 
  • Other benefits as per company practice 

Job Nature: Mon–Fri (Hybrid)

Work Time: 10am-7pm or 6pm-3am (Dhaka Time), Flexible  Location: Gulshan 1, Dhaka 

If you have the skills, experience and drive to excel in this challenging and rewarding role, we would love to hear from you. Apply today and take the next step in your career with us! 

WSD is an employer that values diversity. We highly encourage applications from appropriately qualified and eligible candidates irrespective of age, race, religion, national origin, gender, sexual orientation, gender identity and/or expression, veteran status, disability, or any other status protected by applicable law.