About this role
Hiringhood seeks a Data Scientist to develop and deploy AI/ML solutions for identity management using real-world identity data. This is a full-time, work-from-office position in Hyderabad, India, requiring 7–15 years of experience. The role calls for hands-on Python coding and independent problem-solving; access to external LLM coding assistants may be limited because of data confidentiality.
Responsibilities
- Build name-matching, entity-resolution, record-linkage, duplicate-detection and identity-matching algorithms. Apply NLP and text-similarity techniques such as Jaro-Winkler, Levenshtein, TF-IDF, n-grams, token-based and phonetic matching, embeddings and transformer-based approaches.
- Prepare data, engineer features, conduct statistical analysis, train and validate models, and optimize performance. Combine deterministic rules, fuzzy matching, statistical methods and machine learning in hybrid solutions.
- Write maintainable Python code and reusable ML components. Understand, improve and extend an existing Python codebase, including identifying issues and optimizing algorithms.
- Deploy models as production services or APIs and collaborate with engineering teams on enterprise integration. Assess false positives, false negatives, scalability, latency and explainability. Evaluate AI/NLP techniques for identity verification, fraud detection and identity intelligence.
Qualifications
- Strong Python software-development skills and a foundation in data science, machine learning, statistics, predictive modelling, NLP, text processing, similarity algorithms and entity matching. Experience with Pandas, NumPy, Scikit-learn and SciPy; identity management, digital identity, KYC/e-KYC or biometric applications; customer or entity resolution and duplicate identity detection; multilingual NLP, transliteration and cross-script name matching; and ML deployment, REST APIs, Docker, AWS or MLOps.
- Ability to modify an existing codebase and solve problems with limited reliance on LLM coding assistants. A bachelor's or master's degree in Computer Science, Data Science, Artificial Intelligence, Statistics, Mathematics, Engineering or a related discipline is required.
- Experience with fraud detection and anomaly-detection models for identity fraud is an added advantage.
Skills for this role
PythonData ScienceMachine LearningNatural Language ProcessingText matchingEntity resolutionRecord linkageDuplicate detectionJaro-WinklerLevenshtein distanceTF-IDFN-gramsPhonetic matchingEmbeddingsTransformersFeature engineeringStatistical analysisPredictive modellingPandasNumPyScikit-learnSciPyIdentity managementKYCBiometricsMultilingual NLPTransliterationFraud detectionAnomaly detectionML model deployment גREST APIs Docker AWS MLOps