SPG Consulting

Spark + Scala+ Python +Github Copilot

SPG Consulting Bengaluru, Karnataka, India

Staffing and Recruiting · 11-50 employees

19 h ago
python Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, develop, and maintain scalable data processing pipelines using Apache Spark, Scala, and Python. Leverage AI-assisted tools like GitHub Copilot to enhance code quality, debugging, and development productivity.

What they look for

Apache Spark Scala Python PySpark GitHub Copilot Data Engineering Spark SQL Distributed Computing Git GitHub ETL Data Pipelines Performance Optimization Cloud Platforms Data Structures Algorithms

Requirements

Requires 4-8 years of experience in data engineering with strong proficiency in Apache Spark, Scala, and Python. Candidates must have a solid understanding of distributed computing, data structures, and modern data engineering practices.

Full description

Job Description – Spark + Scala + Python + GitHub Copilot

Position

Senior Data Engineer – Apache Spark / Scala / Python

Experience

4–8 years

Job Summary

We are looking for an experienced Data Engineer with strong expertise in Apache Spark, Scala, Python, and GitHub Copilot. The candidate will be responsible for developing scalable data processing solutions, building data pipelines, optimizing Spark workloads, and leveraging AI-assisted development tools to improve engineering productivity and code quality.

Key Responsibilities

  • Design, develop, and maintain scalable data processing pipelines using Apache Spark.
  • Develop high-performance Spark applications using Scala and Python (PySpark).
  • Build batch and near-real-time data processing solutions.
  • Perform data transformation, cleansing, aggregation, and enrichment using Spark.
  • Optimize Spark jobs for performance, scalability, memory utilization, and cost.
  • Work with large datasets across distributed data platforms.
  • Develop reusable and maintainable Scala/Python code following coding standards.
  • Troubleshoot Spark jobs, performance issues, data-quality problems, and production failures.
  • Implement data validation, error handling, logging, and monitoring.
  • Work with cloud-based data platforms and distributed storage systems.
  • Use Git/GitHub for source control, branching, code reviews, and collaboration.
  • Leverage GitHub Copilot for code generation, refactoring, unit tests, documentation, SQL, and development productivity.
  • Review and validate Copilot-generated code for correctness, security, performance, and maintainability.
  • Collaborate with Data Architects, Data Engineers, Analysts, and business stakeholders.
  • Participate in Agile ceremonies, technical discussions, code reviews, and production support.

Required Skills

  • Strong hands-on experience with Apache Spark.
  • Strong programming experience in Scala.
  • Good hands-on experience with Python/PySpark.
  • Strong understanding of Spark SQL, DataFrames, and RDDs.
  • Experience with distributed computing and large-scale data processing.
  • Strong knowledge of data structures, algorithms, and performance optimization.
  • Experience with Git and GitHub.
  • Hands-on experience with GitHub Copilot or similar AI-assisted development tools.
  • Good understanding of ETL/ELT concepts and data engineering principles.
  • Strong debugging and problem-solving skills.
  • Good communication and collaboration skills.

Good to Have

  • Experience with Azure Databricks, AWS EMR, or Google Cloud Dataproc.
  • Knowledge of Delta Lake / Delta Tables.
  • Experience with Apache Kafka or other streaming technologies.
  • Knowledge of Hive, HDFS, and Hadoop ecosystem.
  • Experience with Azure Data Factory, AWS Glue, or similar orchestration tools.
  • Knowledge of Docker and Kubernetes.
  • Experience with CI/CD and DevOps practices.
  • Knowledge of SQL and relational databases.
  • Experience with cloud data warehouses such as Snowflake, Azure Synapse, or BigQuery.

GitHub Copilot Expectations

  • Use GitHub Copilot to accelerate development of Scala and Python applications.
  • Generate and enhance unit tests, documentation, and repetitive code.
  • Use Copilot for debugging, refactoring, and code optimization.
  • Apply proper engineering judgment when reviewing AI-generated code.
  • Ensure generated code complies with organizational security, coding, and data-governance standards.

Education

Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field.

Key Technologies

Apache Spark | Scala | Python | PySpark | Spark SQL | Git | GitHub | GitHub Copilot | Databricks | Kafka | Hadoop | Cloud

Preferred Candidate Profile

The ideal candidate should have strong hands-on experience in Spark, Scala, and Python, with a solid understanding of distributed data processing and modern data engineering practices. Experience using GitHub Copilot effectively and responsibly to improve development productivity is highly desirable.

Similar roles