About the role
The role involves designing and maintaining scalable data pipelines, developing statistical models, and creating dashboards to provide actionable business insights. Additionally, the candidate will contribute to AI initiatives by developing and monitoring ML and GenAI-based pipelines.
What they look for
Requirements
Candidates must hold a bachelor's degree in a relevant technical field and possess 2-5 years of experience in data engineering or analytics. Essential technical requirements include proficiency in SQL, Python, data modeling, and experience with the Hadoop ecosystem.
Full description
Key Responsibilities:
- Design, implement, and maintain scalable data pipelines for the extraction, transformation, and loading (ETL) of large datasets
- Create algorithms and statistical models to extract actionable insights.
- Design, develop and implement data models for base tables and dashboards
- Conduct thorough analysis to understand data models, upstream source systems, and trends within datasets
- Query massive data sets to interpret complex relations
- Build complex logics for attributes, metrics and feature banks under engineered datasets and reports
- Leverage GenAI capabilities and machine learning techniques to enhance data analysis capabilities.
- Enable and/or develop AI use cases (ML or GenAI based)
- Test, vet, deploy, maintain and monitor AI pipelines
- Create visualizations and dashboards to communicate data-driven insights.
- Package and serve data products over the various reporting outlets in use
- Adhere to information security and personal data privacy mandates and guidelines in data collection, analysis, access management and reporting
- Implement best practices for data governance, quality and documentation
- Implement best practices and continuously improve data analytics processes
- Develop and maintain data architecture and data management, ensuring data integrity, security, and optimal performance
Education:
Bachelor's degree in IT, Computer Engineering, Computer Science, Data Science
Level of Experience:
Limited Experience (2-5Yrs) in a related field
Technical Skills & Knowledge:
Essential:
- Excellent knowledge of data warehousing and data modelling principle
- Excellent knowledge SQL
- Very good knowledge of Linux and OS Administration
- Excellent knowledge of Python
- Working knowledge of orchestration platforms (e.g. Airflow)
Desirable:
- Good knowledge of telecom core systems and data sets
- Good knowledge of key information security and networking principles
- Good knowledge of Scala
- Good knowledge of spark framework
Certifications & Licensure
Essential:
- SQL (any variant) Certification
- Python Certification
- Hadoop or Data Lakehouse Certification
Desirable:
- GenAI and LLM Engineering Certification
- Spark Certification
- Airflow Certification
Tools & Systems:
Essential:
- Data Warehousing or modern data platform
- Apache Hadoop eco-system
- Power BI (or other visualization tools)
- Python (base, pandas, scikit-learn)
- Agentic Development
Desirable:
- GenAI e2e solutions development
Similar roles
-
Sr. Data Analyst
Dynanet Corporation Washington, District of Columbia, United States · $160K–$180K/yr
-
Quality & Compliance Data Analyst - 5828
ColumbiaCare Services Portland, Oregon, United States · $58K–$73K/yr
-
Risk Data Analyst
Oregon Health & Science University Portland, Oregon, United States · $70K–$112K/yr
-
Associate / Staff Mission Data Analyst
SciTec Boulder, Colorado, United States · $88K–$127K/yr
-
Data Analyst
NYU Langone Health New York, New York, United States · $70K–$83K/yr
-
DATA ANALYST
Franklin Precision Industry Inc Franklin, Kentucky, United States