About the role
The role involves administering and optimizing enterprise Big Data platforms and their underlying infrastructure using automation and AI-assisted tools. You will be responsible for managing Linux systems, databases, and containerized environments while ensuring high availability and performance.
What they look for
Requirements
Candidates must have 3-5 years of experience in Linux system administration, database management, or Big Data infrastructure. A bachelor's degree in a technical field and strong proficiency in scripting, automation, and distributed systems are required.
Full description
Job Overview We are seeking a Big Data Infrastructure Engineer with a strong AI-first and automation-driven mindset to join our infrastructure team and support the administration and operation of enterprise Big Data platforms and their underlying infrastructure. The ideal candidate should have strong hands-on experience in Linux operating system administration and database administration, with particular focus on configuration, troubleshooting, performance tuning, optimization, monitoring, and production support. The candidate is expected to actively leverage AI-assisted tools and automation to improve troubleshooting, operational efficiency, scripting, documentation, research, and day-to-day engineering activities, while maintaining the technical judgment required to validate solutions before implementation. The role also requires practical knowledge of Docker and Kubernetes, along with a good understanding of clustering, high availability, distributed systems, and Big Data concepts. A strong willingness to continuously learn and stay current with emerging AI, automation, infrastructure, and Big Data technologies is essential.
Duties and Responsibilities
- Leverage AI-assisted tools to improve troubleshooting, log analysis, scripting, documentation, research, and
operational efficiency while validating outputs before implementation.
- Identify repetitive operational activities and develop automation solutions using Shell/Bash, Python, Ansible,
APIs, or other appropriate technologies.
- Continuously evaluate emerging AI and automation capabilities and identify practical opportunities to improve
infrastructure operations and engineering workflows.
- Administer, configure, manage, troubleshoot, and optimize Linux operating systems supporting production and
non-production environments.
- Monitor and analyze CPU, memory, disk, filesystem, network, processes, and system services, and perform
configuration and performance tuning when required.
- Administer and manage PostgreSQL, MySQL/MariaDB, and Redis, including configuration, access management,
backup and recovery, monitoring, troubleshooting, maintenance, and performance optimization.
- Support database replication, high availability, backup/recovery, and capacity management requirements.
- Support and maintain Docker and Kubernetes environments, including deployment, configuration, monitoring,
troubleshooting, scaling, and cluster administration.
- Support clustered and distributed platforms with focus on high availability, replication, failover, load
balancing, quorum, capacity management, and disaster recovery.
- Support the installation, configuration, monitoring, administration, and upgrade of Cloudera/Hortonworks and
Hadoop-based environments.
- Support and troubleshoot Hadoop ecosystem components such as HDFS, YARN, Hive, Spark, HBase, and
Kafka, as well as related platforms such as Airflow, Superset, and Trino/Presto where applicable.
- Perform production monitoring and support using tools such as Zabbix and Grafana, and participate in incident
management, root cause analysis, and corrective/preventive actions.
- Support security integrations and technologies such as Ranger, LDAP, and Kerberos.
- Collaborate with development, infrastructure, and other technical teams on deployments, upgrades,
infrastructure changes, troubleshooting, and production support.
- Maintain technical documentation, operational procedures, automation, and infrastructure configuration records.
Skills and Qualifications
- 3-5 years of relevant hands-on experience in Linux/System Administration, Database Administration, Big Data
Infrastructure, DevOps, or a related infrastructure role.
- Bachelor's Degree in Computer Science, Computer Engineering, Information Technology, or a related field.
- Strong AI-first and automation-driven mindset, with demonstrated ability to use AI-assisted tools effectively in
technical workflows and critically validate generated recommendations before applying them. Big Data Infrastructure Engineer - Job Description Good scripting and automation skills using Shell/Bash; knowledge of Python, Ansible, APIs, or similar technologies is highly desirable.
- Strong hands-on knowledge of Linux administration, including system configuration, service management,
resource management, storage/filesystems, permissions, networking, troubleshooting, and performance optimization.
- Good hands-on knowledge of PostgreSQL, MySQL/MariaDB, and Redis administration, including configuration,
backup and recovery, users and privileges, monitoring, maintenance, and performance tuning.
- Good understanding of database concepts including connections, transactions, locks, indexing, query
performance, replication, and high availability.
- Good hands-on understanding of Docker and Kubernetes, including containers, images, pods, deployments,
services, storage, networking, monitoring, resource management, and troubleshooting.
- Good understanding of clustering and distributed system concepts, including high availability, replication,
failover, load balancing, and quorum.
- Good understanding of networking fundamentals, including TCP/IP, DNS, ports, routing, connectivity, and
network troubleshooting.
- Good understanding of Big Data concepts and the Hadoop ecosystem, with familiarity or hands-on experience
in HDFS, YARN, Hive, Spark, Kafka, and HBase.
- Familiarity with Cloudera or Hortonworks platforms is highly desirable.
- Familiarity with Zabbix/Grafana, Ranger/LDAP/Kerberos, CI/CD tools, and Trino/Presto is an advantage.
- Strong troubleshooting, analytical, and problem-solving skills, with the ability to investigate issues
systematically and identify root causes.
- Ability to work effectively in production environments, collaborate across technical teams, take ownership of
assigned activities, and continuously develop technical knowledge.
Preferred Certifications
- Relevant Linux certifications such as RHCSA or RHCE.
- Kubernetes certification such as CKA.
- PostgreSQL or MySQL-related certifications/training.
- Red Hat Ansible or other relevant automation certifications.
Note: Certifications are considered an advantage and are not a substitute for practical hands-on experience.