CURIOUS LEARNER DATA SCIENTIST · PYTHON DEVELOPER

Archit Roy

I build production data-science systems where models meet products: pricing, behavioural prediction, generative AI, evaluation, APIs and the machinery that keeps them trustworthy.

I build production data-science systems at the point where models meet products: dynamic pricing, behavioural prediction, generative AI, evaluation, APIs and the operational machinery that keeps them trustworthy.

My default language is Python. I enjoy going from first-principles research and exploratory notebooks to data contracts, failure modes, observability and maintained services. I care as much about proving when a model should not act as I do about improving its score.

Data Scientist 2

Zoomcar

Own and improve production pricing systems across base fares, product and flexibility recommendations, dynamic multipliers, asset controls, alerting and audited business changes.

Data Scientist 1

Zoomcar

Built and operated ML and generative-AI systems for booking intent, marketplace imagery, evidence-grounded support automation and analytical observability; designed experiments and production evaluation around them.

Data Science Intern

Zoomcar

Researched twenty-five years of text-summarisation approaches and built DS-Nemo: an evaluated offline generation pipeline with low-latency APIs serving more than 200K daily requests.

Review summarisation project visual
01 / DS-NEMO

Review summarisation

A review-summarisation system developed through classical NLP, fine-tuned transformers and LLMs, then productionised around evaluation and low-latency serving.

Associated with a 46% reduction in review clicks and a 12% reduction in checkout-session duration. The API handled 200K+ daily requests with an uncached p99 of 17 ms.
GEN AIMLPythonFlask + GunicornGeminiDeepEvalLangfuse
READ CASE STUDY →
Booking-intent ranking project visual
02 / DS-OSIRIS

Booking-intent ranking

A production propensity system and controlled experiments that separated predictive ranking from intervention impact.

The top decile converted at about 3.16× baseline; experiments showed that a strong ranker does not guarantee a useful intervention.
MLPythonXGBoostScikit-learnBigQueryDocker
READ CASE STUDY →
Marketplace pricing operating system project visual
03 / PRICING

Marketplace pricing operating system

Base fares, behavioural ranges, dynamic multipliers, product economics, interventions and change management treated as one ecosystem.

Made high-impact pricing changes faster, more explicit and easier to trace without a disruptive rewrite.
MLPRICINGSTATISTICALPythonSQLKafkaBigQueryMySQL
READ CASE STUDY →
Guarded image enhancement project visual
04 / DS-PIXEL

Guarded image enhancement

An end-to-end generative image pipeline with object-consistency checks, aesthetic scoring and marketplace experiments.

Produced directional CTR and ranking improvements across city experiments while filtering many harmful edits.
GEN AIMLPythonGenerative AICloud VisionCLIP/LAIONFlask
READ CASE STUDY →
Outbound review voice agent project visual
05 / AIB-TANISHA

Outbound review voice agent

A policy-aware multilingual voice campaign for giving the platform's quiet majority of friction-free trips a fair chance to be represented in app-store reviews.

Produced a concise, versioned conversation specification and adversarial scenario matrix for the campaign.
GEN AIVoice AIPrompt designState machinesMultilingual UXEvaluation
READ CASE STUDY →
Evidence-grounded support automation project visual
06 / DS-CIPHER

Evidence-grounded support automation

An LLM-assisted support system improved through manual false-positive analysis, stakeholder adjudication and duplicate-safe processing.

Reached roughly 51% automated resolution across eligible flows at about 94% audited accuracy.
GEN AIPythonGeminiFlaskSQLAlchemyMySQL
READ CASE STUDY →
Pricing anomaly evaluation project visual
07 / DS-WATCHDOG

Pricing anomaly evaluation

A daily evaluator that answers a simple question: did a business metric genuinely move, or are we looking at normal variation or incomplete data?

Converted an exploratory workbook into a replay-tested daily service with explicit data-quality states.
MLSTATISTICALPythonBigQueryMySQLStatisticsData quality
READ CASE STUDY →
Operational reporting platform project visual
08 / DS-PIGEON

Operational reporting platform

A common reporting layer that replaced scattered mailers with gated, previewable and consistent operational communication.

Created a reusable last mile for analytical products without duplicating their business logic.
INFRASTATISTICALPythonHTML emailSMTPMySQLOperational tooling
READ CASE STUDY →
Central DS–DE reliability alerts project visual
09 / DS-ALERTS

Central DS–DE reliability alerts

A central reliability layer for execution health, infrastructure failures and database integrity across production data-science and data-engineering services.

Created a common operational safety net across the DS–DE production estate.
INFRASTATISTICALPythonObservabilityDatabasesInfrastructureAlert routing
READ CASE STUDY →

LANGUAGES

Python · SQL · C++

ML + STATS

XGBoost · Scikit-learn · SHAP · Pandas · NumPy · Experimentation

NLP + GENERATIVE AI

Gemini · Transformers · T5 · BART · PEGASUS · RAG · LLM evaluation

SYSTEMS

Flask · Gunicorn · SQLAlchemy · Docker · Kubernetes · Kafka

DATA + CLOUD

BigQuery · MySQL · GCP · Vertex AI · ETL pipelines

OBSERVABILITY

Langfuse · OpenTelemetry · New Relic · Last9 · Jenkins · ArgoCD

Vellore Institute of Technology

B.Tech, Information Technology · CGPA 9.09 / 10.0

Kunskapsskolan

CBSE · Grade 12: 94% · Grade 10: 91%

AWS Certified Cloud Practitioner · Microsoft Power Platform App Maker Associate