Shubham Kumar Jha

ML & Data Solutions

Building end-to-end machine learning and data solutions — from data preparation and model evaluation to SQL analytics, interactive BI, APIs, and deployment.

Machine Learning

Classification • Model Evaluation • Scikit-Learn

Build, compare and evaluate ML models using structured preprocessing and validation.

Data Analytics

Python • Pandas • SQL • EDA

Transform raw data into insights through cleaning, analysis and exploratory workflows.

Business Intelligence

Power BI • DAX • PostgreSQL

Build interactive dashboards, KPIs and analytical views for business reporting.

End-to-End Delivery

FastAPI • Render • Git • GitHub

Take projects from development through API integration, deployment and version control.

ML / DS Classification & Model Evaluation
ANALYTICS Python • Pandas • SQL
BI Power BI • DAX • PostgreSQL
DEPLOYMENT FastAPI • Render • GitHub
Shubham Kumar Jha
Building reproducible ML workflows. Solving real business problems.

Professional Profile

I am an aspiring Machine Learning Engineer and Data Scientist focused on building end-to-end ML pipelines, SQL analytics layers, and interactive BI dashboards. My work spans data preparation, model evaluation, API development, deployment, and data modeling.

Engineering Philosophy

  • Context First: Understand the problem before choosing a model or analytical approach.
  • Data Quality: Treat preprocessing, validation and data structure as first-class engineering concerns.
  • Empirical Evaluation: Compare models and evaluate them using appropriate metrics instead of assuming one algorithm is best.
  • End-to-End Thinking: Take projects beyond notebooks into APIs, deployment, SQL layers, dashboards and usable interfaces.
  • Reproducibility: Keep workflows structured, versioned and explainable.

Current Technical Focus

Machine Learning Classification Feature Engineering Model Evaluation FastAPI PostgreSQL SQL Power BI DAX DirectQuery Data Visualization

Career Objective

My objective is to build and deploy practical machine learning and data solutions — working across model development, API engineering, SQL analytics, and BI to deliver usable, end-to-end data products.

Project Workflow

Business Problem → Data Preparation → Feature Engineering → Model Comparison → Evaluation → Deployment / Dashboard

Engineering Capabilities

My engineering capabilities span machine learning, SQL data modeling, business intelligence, and application deployment — developed through end-to-end personal projects and industry internship experience.

Machine Learning

Built end-to-end classification pipelines comparing multiple models — selecting Random Forest through empirical evaluation with Scikit-Learn.

Python Scikit-Learn Classification Random Forest Gradient Boosting Feature Engineering Model Evaluation

Data Analytics & SQL

Transforming raw datasets into structured analytical and machine learning workflows through preprocessing, feature engineering, SQL modeling and exploratory analysis.

Python Pandas NumPy SQL PostgreSQL Data Cleaning EDA Data Transformation

BI & Visualization

Built a 9-page Power BI dashboard with 26 DAX measures connected via DirectQuery to a PostgreSQL BI semantic layer.

Power BI DAX DirectQuery Dashboarding KPI Development Data Visualization

Deployment

Built deployable ML applications using FastAPI and Render, with additional experience using Streamlit and Flask for application interfaces.

FastAPI Render Flask Streamlit Git GitHub

Featured Projects

Room Type Predictor UI
ML Classification • API • Deployment

Room Type Predictor

NYC Airbnb Room Type Classification
Python Pandas Scikit-Learn Random Forest FastAPI Render
Accuracy: 85.6%  ·  Macro F1: 0.742
View Full Case Study Live Demo
Olist Power BI Dashboard
BI Dashboard • SQL • DAX

Olist E-Commerce BI Dashboard

End-to-End E-Commerce Analytics Pipeline
PostgreSQL SQL Power BI DAX DirectQuery
Revenue: $14.21M
Orders: 98.67K
On-time: 93.23%
Avg Rating: 4.03
View Full Case Study View Dashboard Screenshots

Experience

Jan 2026 – Jun 2026

Data Science Trainee

Data Science

Solitaire Infosys Pvt. Ltd. | Mohali, Punjab

  • Engineered data cleaning pipelines for a telecom client's dataset, resulting in structured and query-ready data for downstream modeling.
  • Evaluated dataset schema to isolate key behavioral features, successfully establishing the foundation for churn prediction analysis.
  • Collaborated within a cross-functional data team to apply enterprise data handling standards, ensuring consistent and reproducible analytic workflows.
Data Cleaning EDA Client Data Team Collaboration
Jan 2025 – Mar 2025

AI Prompt Evaluator — NLP Model QC

AI / RLHF

Data Annotation | Remote (California, UK & EU)

  • Evaluated language model completions across complex reasoning tasks, identifying logical inconsistencies to improve overall model safety.
  • Optimized AI alignment pipelines by diagnosing contextual hallucinations, producing structured feedback that directly informed model refinement.
  • Engineered evaluation rubrics for edge-case scenarios, ensuring high consistency in human-feedback data provided to training teams.
LLM Evaluation RLHF AI Alignment NLP Remote
Apr 2024 – Sep 2024

AI Trainer — Prompt Engineering & Evaluation

Prompt Engineering

Outlier | Remote (USA)

  • Developed targeted prompt variations for multi-turn conversational tasks, empirically improving model adherence to complex NLP instructions.
  • Evaluated generative outputs across technical domains, identifying logic errors to isolate specific vulnerabilities in model reasoning.
  • Designed structured feedback loops for agile data teams, accelerating the refinement of high-quality training datasets for AI optimization.
Prompt Engineering LLM Evaluation NLP AI Model Tuning USA Remote

Certifications

freeCodeCamp
Data Analysis with Python
July 2025
freeCodeCamp
Machine Learning with Python
July 2025

Achievements

Key milestones reflecting technical growth, engineering experience, industry exposure, and professional recognition in Machine Learning and Data Science.

1st Place
Project Innovation Competition
Won 1st place among 1,000+ participants by presenting a Machine Learning project demonstrating end-to-end model development and deployment.
Global
International AI Project Experience
Collaborated remotely with international AI teams through Outlier, contributing to LLM evaluation, prompt quality improvement, annotation workflows, and model response assessment.
85.6%
Best ML Classification Accuracy
Designed, tuned, and deployed a complete Random Forest classification application that achieved 85.6% test accuracy and 0.742 Macro F1 on the NYC Airbnb dataset.

Education

2022 – 2026 (Expected)

B.Tech — CSE (Data Science)

Gulzar Group of Institutes, PTU | Ludhiana, Punjab

Data Science Machine Learning Statistics DBMS Python Programming
2021 – 2022

Class XII — Science (PCM)

CBSE Board

2019 – 2020

Class X

CBSE Board

If you're building something with data, let's talk.

Open to internship and entry-level opportunities in Machine Learning and Data Science.

▸ ML Engineer  ·  Data Scientist  ·  AI/NLP Roles

+91-8448517087
India — Open to Remote