Senior Data Scientist
INFO ORIGIN INC Woodlawn, Maryland, United States · $166K–$187K/yr
Software Development · 201-500 employees
About the role
The Senior Data Scientist will design and maintain advanced data processing and entity resolution pipelines using Python and SQL. They will also collaborate with stakeholders to implement NLP and data matching techniques to improve data quality across large-scale datasets.
What they look for
Requirements
Candidates must have a Bachelor's or Master's degree in a relevant field and at least 10 years of IT experience. Strong hands-on proficiency in Python, SQL, NLP, and entity resolution techniques is required.
Full description
Senior Data Scientist
Location: Woodlawn, MD Job Type: Long-Term Contract Work Arrangement: Onsite Interview: Video Interview Duration: Long Term
Job Overview
We are seeking an experienced Senior Data Scientist to join our team in Woodlawn, MD. The ideal candidate will have strong hands-on experience in Natural Language Processing (NLP), text processing, information extraction, Python, SQL, and data matching/entity resolution.
This role will focus on developing advanced data processing solutions, building scalable data pipelines, improving data quality, and implementing techniques to identify and match records across large and complex datasets.
Key Responsibilities
- Design, develop, and maintain advanced data processing and entity resolution pipelines using Python and SQL.
- Process, clean, transform, and manage large-scale datasets from multiple sources.
- Apply NLP, text processing, and information extraction techniques to structured and unstructured data.
- Implement data matching techniques such as Named Entity Recognition (NER), blocking and indexing, string similarity/distance metrics, TF-IDF, cosine similarity, phonetic matching, and address standardization.
- Develop data cleansing and validation solutions using Python and Regex.
- Write and optimize complex SQL queries and database operations for performance and scalability.
- Perform data validation, testing, deployment, and production monitoring.
- Participate in code reviews and follow version control, data security, reproducibility, and software development best practices.
- Collaborate with technical teams and business stakeholders to understand requirements and translate complex data-processing logic into practical solutions.
- Troubleshoot data quality and pipeline issues and implement reliable, scalable solutions.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Statistics, Applied Mathematics, Information Science, or a related field.
- 10+ years of IT experience.
- Strong hands-on experience with Python and SQL.
- Strong practical experience with NLP, text processing, and information extraction.
- Experience with Named Entity Recognition (NER).
- Experience with entity resolution, record linkage, data matching, or deduplication.
- Knowledge of blocking/indexing and string distance metrics.
- Experience with TF-IDF and cosine similarity.
- Experience with Regex for text processing and data cleansing.
- Knowledge of address standardization and phonetic encoding.
- Strong analytical, problem-solving, written, and verbal communication skills.
Preferred Skills
Experience with one or more of the following is highly desirable:
- spaCy
- Scikit-learn
- Splink
- Dedupe
- FastLink
- recordlinkage
- PostgreSQL
- DB2
- Oracle
- SQL Server
- Hadoop
- Jenkins
- CI/CD
- Apache Airflow or other pipeline automation tools
Preferred Experience
- Experience working on federal, state, or local government projects.
- Experience working with large-scale, legacy, or distributed data systems.
- Experience migrating and integrating data from multiple enterprise sources.
- Experience designing automated data-quality and validation pipelines.
- Ability to explain complex data-matching algorithms and business rules to technical and non-technical stakeholders.
- Ability to independently own projects from data discovery through development, deployment, and post-production support.
Similar roles
-
[8BE] Data Scientist (AI + ML)
Software Mind Buenos Aires, Argentina
-
Staff Data Scientist, Optimization
Firestorm San Diego, California, United States · $175K–$220K/yr
-
Lead Data Scientist
HUCKEYE HEALTH SERVICES LLC Houston, Texas, United States · $106K–$135K/yr
-
Principal Data Scientist, Google Search
Google Mountain View, California, United States · $307K–$427K/yr
-
Junior Data Scientist
Skidmore College Saratoga Springs, New York, United States
-
Data Scientist II
CorVel Corporation Irvine, California, United States · $83K–$127K/yr