Principal Data Engineer - AWS
Metis Technology Solutions Inc Santa Clara County, California, United States · $140K–$200K/yr
Aviation and Aerospace Component Manufacturing · 51-200 employees
About the role
The Principal Data Engineer will architect, develop, and maintain scalable ETL/ELT data pipelines and cloud infrastructure within the AWS ecosystem. They will collaborate with researchers and developers to design automated data ingestion solutions and ensure the reliability, security, and performance of data-processing systems.
What they look for
Requirements
Candidates must possess at least 10 years of relevant technical experience, including 5 years of hands-on AWS experience in production environments. A bachelor's degree in a technical discipline is required, along with strong proficiency in Python, SQL, and infrastructure-as-code tools like Terraform.
Full description
POSITION SUMMARY:
Metis Technology Solutions is seeking an experienced Principal Data Engineer – AWS to join our Software Operations team. This engineer will work closely with software developers, researchers, analysts, database engineers, and system administrators to design, develop, deploy, and operate the data infrastructure and pipelines connecting research platforms, external data sources, cloud-based services, and the project's Sherlock data warehouse.
This position:
- Architects, designs, develops, deploys, and maintains scalable and reliable ETL/ELT data pipelines in Amazon Web Services (AWS).
- Designs automated ingestion and processing solutions for structured, semi-structured, unstructured, batch, and streaming data.
- Develops production-quality data-processing and pipeline software using Python, SQL, and shell scripting.
- Designs and maintain workflow orchestration solutions using technologies such as Apache Airflow, Dagster, AWS Step Functions, or equivalent platforms.
- Designs and implements real-time and near-real-time streaming and event-driven data pipelines using technologies such as AWS Kinesis, Amazon Data Firehose, Kafka, RabbitMQ, SQS/SNS, or equivalent technologies.
- Develops and maintains cloud data solutions using AWS services such as Amazon S3, AWS Lambda, Amazon Redshift, Amazon RDS, Amazon DynamoDB, Amazon EC2, API Gateway, IAM, and CloudWatch, as appropriate to project requirements.
- Administers and optimizes relational databases and cloud data warehouses, including PostgreSQL/PostGIS and Amazon Redshift or comparable technologies.
- Develops and maintains infrastructure as code (IaC) using Terraform or equivalent technologies.
- Develops and maintains automated deployment and CI/CD processes for data applications and supporting infrastructure.
- Designs pipelines and supporting infrastructure for reliability, scalability, maintainability, observability, security, and efficient use of AWS resources.
- Implements appropriate AWS security practices, including IAM roles and policies, encryption, secrets management, network controls, logging, and least-privilege access.
- Develops monitoring, logging, metrics, and alerting that provide operational visibility into data-pipeline health and performance.
- Documents data architectures, data flows, interfaces, infrastructure, deployment processes, operational procedures, and troubleshooting practices.
- Collaborates with researchers, analysts, and application developers to translate research and application requirements into reliable and maintainable data-processing solutions.
MINIMUM QUALIFICATIONS:
Education:
Bachelor’s degree or higher in computer science, computer engineering, information systems, or a related technical discipline.
Required Skills and knowledge:
- Minimum 10 years of progressively responsible software engineering, data engineering, database engineering, or closely related technical experience.
- Minimum 5 years of substantial hands-on AWS experience, including the design, implementation, deployment, and operation of production data-processing or data-pipeline solutions.
- Demonstrated experience architecting and developing ETL/ELT pipelines involving large, heterogeneous, or rapidly changing datasets.
- Strong hands-on experience with AWS data and compute services. Relevant technologies may include S3, Lambda, Redshift, RDS, DynamoDB, Kinesis, Data Firehose, EC2, API Gateway, IAM, and CloudWatch.
- Advanced SQL skills and substantial experience working with relational database systems, preferably PostgreSQL/PostGIS.
- Demonstrated experience with cloud data warehouses such as Amazon Redshift or comparable technologies.
- Experience with NoSQL databases or document/key-value data stores such as DynamoDB, MongoDB, or comparable technologies.
- Demonstrated experience with data-pipeline and workflow orchestration using Apache Airflow, Dagster, AWS Step Functions, or comparable technologies.
- Experience designing or implementing streaming, message-oriented, or event-driven data-processing systems using technologies such as Kinesis, Data Firehose, Kafka, RabbitMQ, SQS/SNS, or comparable technologies.
- Demonstrated experience diagnosing and correcting database and data-pipeline performance problems.
- Hands-on experience with infrastructure as code, preferably Terraform or an equivalent technology.
- Working knowledge of AWS security concepts including IAM, encryption, secrets management, network security, logging, and least-privilege access.
- Strong understanding of modern software-engineering practices, including modular design, automated testing, version control, documentation, security, and maintainable code.
- Strong analytical and troubleshooting skills and demonstrated ability to diagnose complex problems spanning applications, databases, data pipelines, and cloud infrastructure.
- Proven ability to evaluate technical alternatives, make sound architectural decisions, and communicate the rationale and tradeoffs associated with those decisions.
- Excellent written and verbal communication skills and demonstrated ability to collaborate effectively with software engineers, researchers, analysts, system administrators, and other technical stakeholders.
LOCATION SPECIFIC REQUIRMENTS: NASA Ames Research Center, Moffett Field CA
SECURITY CLEARANCE: Applicant must be eligible to obtain a U.S. Government Public Trust Clearance. Must be a U.S. Citizen or Permanent Resident.
EEOE Including Vets and Disability
Similar roles
-
Data Engineer - Splunk and Azure
Bosch Group Bengaluru, Karnataka, India
-
Data Engineer
BID Operations Shenzhen, Guangdong Province, China
-
Data Engineer
Prodigal Bengaluru, Karnataka, India
-
Lead Data Engineer
Bristlecone Noida, Uttar Pradesh, India
-
Quantitative Data Engineer
Qube Research & Technologies Hong Kong, Hong Kong Island, Hong Kong S.A.R.
-
Data Engineer
CipherHealth United States