AI/ML Data Engineer (TS/SCI with Poly Required)
GCI Incorporated · Chantilly, Virginia, United States · $164K–$275K/yr
IT Services and IT Consulting · 501-1,000 employees
About the role
The Data Engineer will develop, optimize, and deploy multilingual machine translation models and scalable MLOps pipelines. They will work closely with Data Scientists to manage large datasets and support the training of custom models for various language types.
What they look for
Requirements
Candidates must have a Bachelor's degree in a technical field and at least 10 years of software engineering experience. A US citizenship and an active TS/SCI with Polygraph clearance are strictly required.
Full description
GCI embodies excellence, integrity and professionalism. The employees supporting our customers deliver unique, high-value mission solutions while effectively leverage the technological expertise of our valued workforce to meet critical mission requirements in the areas of Data Analytics and Software Development, Engineering, Targeting and Analysis, Operations, Training, and Cyber Operations. We maximize opportunities for success by building and maintaining trusted and reliable partnerships with our customers and industry.
At GCI, we solve the hard problems. As an AI/ML Data Engineer, a typical day will include the following duties:
JOB DESCRIPTION
The Data Engineer will work closely with the team to advance Human Language Technologies (HLT), with a focus on refining text and audio translation capabilities. This position will be responsible for establishing an environment to train custom models for the top language types currently available in our triage tool. Additionally, this role will enable the organization to develop the capacity to train low-resource languages which may emerge as future priorities. The individual in this role will work closely with Data Scientists, providing comprehensive support and ensuring seamless coverage. A professional who utilizes statistical analysis, programming skills, and machine learning techniques to collect, clean, analyze, and interpret large datasets, extracting valuable insights and creating predictive models.
KEY RESPONSIBILITIES
- Develop, fine-tune, evaluate, and optimize multilingual machine translation models (e.g., NLLB, Opus-MT, MarianMT) to improve translation quality for low-resource languages.
- Build, preprocess, and manage multilingual text and speech datasets for model training, evaluation, and continuous improvement.
- Design, develop, and maintain scalable data pipelines and end-to-end MLOps workflows for data ingestion, model training, deployment, monitoring, and lifecycle management.
- Develop and deploy cloud-native machine learning solutions using AWS services such as SageMaker, Step Functions, and Bedrock.
- Deploy and support machine learning models in production using containerized environments and CI/CD best practices.
- Engineer features and optimize datasets to improve machine learning model performance.
- Conduct testing, validation, benchmarking, and troubleshooting of machine learning models and data pipelines.
- Research and evaluate emerging AI, machine learning, NLP, and speech technologies for mission applications.
EDUCATION AND EXPERIENCE
- Bachelor’s Degree in Computer Science, Electrical or Computer Engineering or a related technical discipline, or the equivalent combination of education, technical training, or work/military experience
- 10+ years of related software engineering experience.
REQUIRED QUALIFICATIONS
- Experience developing software applications using Python.
- Experience training, fine-tuning, evaluating, and optimizing machine learning and deep learning models using modern frameworks and best practices.
- Familiarity with DevOps and MLOps principles, including CI/CD, infrastructure automation, model lifecycle management, monitoring, version control, and software delivery best practices.
- Hands-on experience with AWS cloud services and AI/ML offerings, including S3, EC2, IAM, VPC, SageMaker, Bedrock, Lambda, and related services.
- Experience developing and deploying containerized applications using Docker.
DESIRED QUALIFICATIONS
- Experience fine-tuning transformer-based language translation models (e.g., NLLB, Opus-MT, MarianMT) and working with Hugging Face Transformers.
- Familiarity with experiment tracking, model registries, and dataset versioning tools such as MLflow, Weights & Biases, or DVC.
- Experience with distributed training frameworks such as Ray or PyTorch Distributed
- Hands-on experience with machine learning frameworks such as PyTorch or TensorFlow.
- Experience designing, implementing, and maintaining production-grade MLOps pipelines and automated machine learning workflows supporting model training, deployment, monitoring, and lifecycle management using technologies such as AWS Step Functions, Apache NiFi, Apache Airflow, or similar orchestration platforms.
- Experience developing multilingual NLP, speech processing, or language translation solutions.
- Familiarity with large language models (LLMs), transformer architectures, and generative AI technologies.
- Familiarity with NLP frameworks such as spaCy, NLTK, Stanford CoreNLP, or similar libraries.
- Strong analytical, problem-solving, and communication skills.
- Ability to work independently and collaboratively in a multidisciplinary environment.
- Familiarity with audio processing frameworks such as Librosa, PyAudioAnalysis, OpenSMILE, or similar technologies.
*A candidate must be a US Citizen and requires an active/current TS/SCI with Polygraph clearance.
Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans