cargill

Data Engineer

cargill · DeKalb County, Georgia, United States

Food and Beverage Manufacturing · 10,001+ employees

4 d ago
Mid (2-5 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves designing, building, and maintaining scalable data systems and pipelines to facilitate efficient data processing and decision-making. You will collaborate with cross-functional teams to implement data infrastructure, automated deployment pipelines, and robust data modeling solutions.

What they look for

Data Engineering Cloud Platforms Data Pipelines SQL Python Java Scala Data Modeling Data Architecture Data Governance DevOps Spark Kafka AWS Glue Airflow CI/CD

Requirements

Candidates must have at least 2 years of relevant work experience, with 3 years or more typically preferred. Proficiency in cloud platforms, SQL, data transformation tools, and modern data architecture concepts is required.

Full description

Cargill is committed to providing food and agricultural solutions to nourish the world in a safe, responsible, and sustainable way. Sitting at the heart of the supply chain, we partner with farmers and customers to source, make and deliver products that are vital for living. Our 155,000 team members innovate with purpose, providing customers with life’s essentials so businesses can grow, communities prosper, and consumers live well. With over 160 years of experience as a family company, we look ahead while remaining true to our values. We put people first. We reach higher. We do the right thing—today and for generations to come.

Job Purpose and Impact

The Professional, Data Engineering job designs, builds and maintains moderately complex data systems that enable data analysis and reporting. With limited supervision, this job collaborates to ensure that large sets of data are efficiently processed and made accessible for decision making.

Key Accountabilities

  • DATA & ANALYTICAL SOLUTIONS: Develops moderately complex data products and solutions using advanced data engineering and cloud based technologies, ensuring they are designed and built to be scalable, sustainable and robust.
  • DATA PIPELINES: Maintains and supports the development of streaming and batch data pipelines that facilitate the seamless ingestion of data from various data sources, transform the data into information and move to data stores like data lake, data warehouse and others.
  • DATA SYSTEMS: Reviews existing data systems and architectures to implement the identified areas for improvement and optimization.
  • DATA INFRASTRUCTURE: Helps prepare data infrastructure to support the efficient storage and retrieval of data.
  • DATA FORMATS: Implements appropriate data formats to improve data usability and accessibility across the organization.
  • STAKEHOLDER MANAGEMENT: Partners with multi-functional data and advanced analytic teams to collect requirements and ensure that data solutions meet the functional and non-functional needs of various partners.
  • DATA FRAMEWORKS: Builds moderately complex prototypes to test new concepts and implements data engineering frameworks and architectures to support the improvement of data processing capabilities and advanced analytics initiatives.
  • AUTOMATED DEPLOYMENT PIPELINES: Implements automated deployment pipelines to support improving efficiency of code deployments with fit for purpose governance.
  • DATA MODELING: Performs moderately complex data modeling aligned with the datastore technology to ensure sustainable performance and accessibility.

Qualifications

Minimum requirement of 2 years of relevant work experience. Typically reflects 3 years or more of relevant experience.

Preferred Qualifications

  • CLOUD ENVIRONMENTS: Familiarity with major cloud platforms (AWS, GCP, Azure).
  • DATA ARCHITECTURE: Experience with modern data architectures, including data lakes, data lakehouses, and data hubs, along with related capabilities such as ingestion, governance, modeling, and observability.
  • DATA INGESTION: Proficiency in data collection, ingestion tools (Kafka, AWS Glue), and storage formats (Iceberg, Parquet).
  • DATA STREAMING: Knowledge of streaming architectures and tools (Kafka, Flink).
  • DATA MODELING: Strong background in data transformation and modeling using SQL-based frameworks and orchestration tools (dbt, AWS Glue, Airflow). Experience with modeling concepts like SCD and schema evolution.
  • DATA TRANSFORMATION: Familiarity with using Spark for data transformation, including streaming, performance tuning, and debugging with Spark UI.
  • PROGRAMMING: Proficient with programming in Python, Java, Scala, or similar languages. Expert-level proficiency in SQL for data manipulation and optimization.
  • DEVOPS: Demonstrated experience in DevOps practices, including code management, CI/CD, and deployment strategies.
  • DATA GOVERNANCE: Understanding of data governance principles, including data quality, privacy, and security considerations for data product development and consumption.

The business will not sponsor applicants for work visa for this position.

Equal Opportunity Employer, including Disability/Vet.