Spark + Scala+ Python +Github Copilot
SPG Consulting Bengaluru, Karnataka, India
Staffing and Recruiting · 11-50 employees
About the role
Design, develop, and maintain scalable data processing pipelines using Apache Spark, Scala, and Python. Leverage AI-assisted tools like GitHub Copilot to enhance code quality, debugging, and development productivity.
What they look for
Requirements
Requires 4-8 years of experience in data engineering with strong proficiency in Apache Spark, Scala, and Python. Candidates must have a solid understanding of distributed computing, data structures, and modern data engineering practices.
Full description
Job Description – Spark + Scala + Python + GitHub Copilot
Position
Senior Data Engineer – Apache Spark / Scala / Python
Experience
4–8 years
Job Summary
We are looking for an experienced Data Engineer with strong expertise in Apache Spark, Scala, Python, and GitHub Copilot. The candidate will be responsible for developing scalable data processing solutions, building data pipelines, optimizing Spark workloads, and leveraging AI-assisted development tools to improve engineering productivity and code quality.
Key Responsibilities
- Design, develop, and maintain scalable data processing pipelines using Apache Spark.
- Develop high-performance Spark applications using Scala and Python (PySpark).
- Build batch and near-real-time data processing solutions.
- Perform data transformation, cleansing, aggregation, and enrichment using Spark.
- Optimize Spark jobs for performance, scalability, memory utilization, and cost.
- Work with large datasets across distributed data platforms.
- Develop reusable and maintainable Scala/Python code following coding standards.
- Troubleshoot Spark jobs, performance issues, data-quality problems, and production failures.
- Implement data validation, error handling, logging, and monitoring.
- Work with cloud-based data platforms and distributed storage systems.
- Use Git/GitHub for source control, branching, code reviews, and collaboration.
- Leverage GitHub Copilot for code generation, refactoring, unit tests, documentation, SQL, and development productivity.
- Review and validate Copilot-generated code for correctness, security, performance, and maintainability.
- Collaborate with Data Architects, Data Engineers, Analysts, and business stakeholders.
- Participate in Agile ceremonies, technical discussions, code reviews, and production support.
Required Skills
- Strong hands-on experience with Apache Spark.
- Strong programming experience in Scala.
- Good hands-on experience with Python/PySpark.
- Strong understanding of Spark SQL, DataFrames, and RDDs.
- Experience with distributed computing and large-scale data processing.
- Strong knowledge of data structures, algorithms, and performance optimization.
- Experience with Git and GitHub.
- Hands-on experience with GitHub Copilot or similar AI-assisted development tools.
- Good understanding of ETL/ELT concepts and data engineering principles.
- Strong debugging and problem-solving skills.
- Good communication and collaboration skills.
Good to Have
- Experience with Azure Databricks, AWS EMR, or Google Cloud Dataproc.
- Knowledge of Delta Lake / Delta Tables.
- Experience with Apache Kafka or other streaming technologies.
- Knowledge of Hive, HDFS, and Hadoop ecosystem.
- Experience with Azure Data Factory, AWS Glue, or similar orchestration tools.
- Knowledge of Docker and Kubernetes.
- Experience with CI/CD and DevOps practices.
- Knowledge of SQL and relational databases.
- Experience with cloud data warehouses such as Snowflake, Azure Synapse, or BigQuery.
GitHub Copilot Expectations
- Use GitHub Copilot to accelerate development of Scala and Python applications.
- Generate and enhance unit tests, documentation, and repetitive code.
- Use Copilot for debugging, refactoring, and code optimization.
- Apply proper engineering judgment when reviewing AI-generated code.
- Ensure generated code complies with organizational security, coding, and data-governance standards.
Education
Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field.
Key Technologies
Apache Spark | Scala | Python | PySpark | Spark SQL | Git | GitHub | GitHub Copilot | Databricks | Kafka | Hadoop | Cloud
Preferred Candidate Profile
The ideal candidate should have strong hands-on experience in Spark, Scala, and Python, with a solid understanding of distributed data processing and modern data engineering practices. Experience using GitHub Copilot effectively and responsibly to improve development productivity is highly desirable.
Similar roles
-
Embedded Python Data & Automation Engineer
Pavago Mexico
-
4549- Software Development Engineer-II (Backend-Python)
Innovaccer Analytics Noida, Uttar Pradesh, India
-
Applied AI Engineer - Python
JoVE Bengaluru, Karnataka, India
-
Senior Backend Engineer (Python), Agent Developer: Flow Components
GitLab United Kingdom
-
Développeur Python / GCP / MS Fabric H/F
NEXTON Saint-Augustin, Nouvelle-Aquitaine, France
-
Senior Full Stack Developer (Python / React)
Jahnel Group Modena, Emilia-Romagna, Italy