Staff Databricks Engineer, Site Reliability Engineering
Kinaxis Inc. Sholinganallur, Tamil Nadu, India
Software Development · 1,001-5,000 employees
About the role
The role involves designing, implementing, and managing complex cloud infrastructure projects while balancing SRE duties with data analytics tasks. You will be responsible for transitioning legacy systems to a modern cloud-native stack and ensuring high service availability through automation and monitoring.
What they look for
Requirements
Candidates must have a bachelor's or master's degree in Computer Science or a related field and over 10 years of experience in site reliability engineering and distributed systems. Deep hands-on expertise with cloud platforms and data tools like Databricks, Snowflake, and Informatica is required.
Benefits
Full description
About Kinaxis
Are you looking to join an innovative, market-leading company where you can truly elevate your career? At Kinaxis we are serious about culture, we are serious about technology, we are serious about customers, and we are serious about not taking ourselves too seriously. If you are looking to be part of an incredible growth story, then we might just be the place for you!
In 1984, we started out as a team of three engineers. Today, we have grown to become a global organization with over 2000 employees around the world, 6 global office and a best-in-class HQ in Ottawa, Canada. As winners of several Top Employer awards globally, we are proud to work with our customers and employees towards solving some of the biggest challenges facing supply chains today.
Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains, and supporting the people who manage them. Our powerful, AI infused platform provides full transparency and visibility across end-to-end supply chains, enabling our customers to make faster, better decisions. We are trusted by renowned global brands to provide the agility and predictability needed to navigate today’s volatility and disruption. With more than 40,000 users in over 100 countries, we are expanding our team as we continue to innovate and revolutionize how we support our customers.
Location
Chennai, India
About the Team
The Site Reliability Engineering team is responsible for the delivery, management, and monitoring of Kinaxis products and cloud infrastructure in our production offerings. We are responsible for ensuring service availability and performance to our customers globally, 7x24x365.
About the Role
The Staff Cloud Engineer, as a seasoned professional with extensive experience, is responsible for the design, implementation, operations, monitoring, automation, testing, and delivery of complex cloud infrastructure projects.
The incumbent will provide technical guidance and expertise as they support the cloud infrastructure solutions delivered in production environments and will act as a lead on interactions with Global Customer Care teams in management of critical escalations.
This role will be 50% SRE & 50% Data Analytics
What you will do
- Transition legacy systems to a cloud-native, modern data stack leveraging tools such as Informatica, Airflow, Postgres and modern technologies like Snowflake, dbt, BigQuery, Looker, CI/CD, Git, Databricks, PowerBI, Datadog and Grafana.
- Deploy, upgrade and support numerous applications, services and operating systems.
- Apply software engineering principles to operational challenges with a focus on automation, self-healing and monitoring solutions.
- Deliver customer excellence, making sure we meet all SLAs.
- Translate architectural, technical and feature changes into operational plans.
- Advise and maintain a feedback loop with development and operations teams to ensure deliveries are in accordance with needs and expectations.
- Deliver cloud reference architectures to be used by teams across the enterprise as they work to modernize existing technologies and build new capabilities to deliver to Customers.
- Develop and manage cloud automation using cloud orchestration capabilities, scripting languages and APIs to design, code, test, implement and support production services.
- Participate in an on-call rotation to investigate incidents and provide root causes relating to production infrastructure, services and applications.
- Troubleshoot issues within the cloud platform, documenting resolutions and providing coaching and feedback to more junior employees on those resolutions
- Responsible for designing, implementing, and managing reporting solutions that support the cloud infrastructure; Optimize cloud resources to ensure cost effectiveness, and identify opportunities for cost savings in customer-facing infrastructure.
- Stay current with emerging cloud technologies and trends, evaluating and recommending new tools and services to enhance our cloud infrastructure.
- Support scalable and secure cloud environments using best practices and industry standards.
What we are looking for
- Bachelor’s/Master's degree in Engineering with specialization in Computer Science or related discipline or demonstrated equivalent experience.
- Prior experience in a site reliability engineering role.
- Deep, hands-on experience with tools such as Snowflake, Databricks, Informatica, dbt, BigQuery, Looker and/or Power BI.
- 10+ years of experience with deployment and monitoring of distributed systems , infrastructure and deployment of Cloud Services, public cloud platforms (both console and API) like GCP, Azure or AWS.
- Working with VMware ESXi is considered an asset.
- Familiarity with cloud infrastructure and supportive Observability tooling (Logstash, Datadog, Grafana).
- Strong knowledge of system design to manage operational and reliability trade-offs.
- Extensive experience developing in a scription language (Ansible, PowerShell, Bash and Python).
- Proven practical experience in building and managing:
- Infrastructure as Code (Terraform)
- Configuration management tools (Ansible)
- CI/CD solutions (Git, GitOps, Argo CD)
- Containers and orchestrators (Docker, Kubernetes, Helm)
- System monitoring and centralized logging platforms (Datadog, Prometheus, ELK)
- In-depth and proactive communication and documentation skills.
#Senior #LI-RJ1 #Fulltime
Work With Impact: Our platform directly helps companies power the world’s supply chains. We see the results of what we do out in the world every day, when we see store shelves stocked, when medications are available for our loved ones, and so much more.
Work with Fortune 500 Brands: Companies across industries trust us to help them take control of their integrated business planning and digital supply chain. Some of our customers include Lockheed Martin, Unilever, P&G, ExxonMobil, Cisco and more.
Social Responsibility at Kinaxis: Our Diversity, Equity, and Inclusion Committee weighs in on hiring practices, talent assessment training materials, and mandatory training on unconscious bias and inclusion fundamentals. Sustainability is key to what we do and we’re committed to a long-term net-zero operations strategy. We are involved in our communities and support causes where we can make the most impact.
People matter at Kinaxis and here are some of the perks and benefits we offer, which may vary by location and employee:
- Flexible vacation and Kinaxis Days (company-wide days off)
- Flexible work options
- Physical and mental well-being programs
- Regularly scheduled virtual fitness classes
- Mentorship programs, training, and career development
- Recognition programs and referral rewards
- Hackathons
For more information, visit the Kinaxis website at www.kinaxis.com or the company’s blog at http://blog.kinaxis.com.
Kinaxis welcomes candidates to apply to our inclusive community. We provide accommodations upon request to ensure fairness and accessibility throughout our recruitment process for all candidates, including those with specific needs or disabilities. If you require an accommodation, please reach out to us at recruitmentprograms@kinaxis.com. This contact information is for accessibility requests only and cannot be used to inquire about the status of applications.
Kinaxis is committed to ensuring a fair and transparent recruitment process. We use artificial intelligence (AI) tools in the initial step of the recruitment process to compare submitted resumes against the job description to identify candidates whose education, experience, and skills most closely match the requirements of the role. After the initial screening, all subsequent decisions regarding your application, including final selection, are made by our human recruitment team. AI does not make any final hiring decisions.
Similar roles
-
Senior Site Reliability Engineer I
Axon Seattle, Washington, United States · $134K–$215K/yr
-
Senior Site Reliability Engineer
Akamai Cambridge, Massachusetts, United States · $186K–$219K/yr
-
Site Reliability Engineer II
Akamai United States · $138K–$171K/yr
-
Analista de SRE Pleno - Vaga Afirmativa para Mulheres
Experian Sao Carlos, Southeast, Brazil
-
Site Reliability Engineer (SRE) Manager
Plume Ljubljana, Slovenia
-
Senior Site Reliability Engineer (SRE)
Tubi - Canada Toronto, Ontario, Canada · CA$116K–CA$235K/yr