Open to AI / Computer Vision roles — remote or relocation

Saurabh Kumar Singh

AI/ML Engineer

Computer Vision · Vision-Language Models · Generative AI

AI/ML Engineer (M.Sc. AI & ML, IIIT Lucknow) with 2+ years of hands-on work in computer vision, generative AI and multimodal model evaluation; currently evaluating Vision-Language Models at Snorkel AI. Built a monocular-depth 2D-to-3D video pipeline, a generative textile-design system and an OCR invoice-extraction pipeline, backed by a B.Sc. in Mathematics.

Years in AI/ML
2+
M.Sc. CGPA / 10
8.00
Professional roles
5
Evaluation focus
VLM

01 — About

What I work on

Vision & multimodal

Detection, segmentation, monocular depth and stereo synthesis — plus evaluating Vision-Language Models on image and video reasoning, where I hunt hallucinations and spatial-reasoning failures and write the corrections that fix them.

Generative AI & evaluation data

Diffusion and image-generation workflows shipped as product features, and expert-level rubrics, reference solutions and preference data for training and benchmarking LLMs and agents.

02 — Toolkit

Technical skills

Languages

  • Python
  • SQL
  • C
  • C++
  • JavaScript

ML / Deep Learning

  • PyTorch
  • TensorFlow
  • Keras
  • Hugging Face Transformers
  • Scikit-learn
  • CNNs
  • Transformers
  • Transfer Learning
  • Hyperparameter Tuning
  • Model Evaluation & Error Analysis

Computer Vision

  • OpenCV
  • Object Detection
  • Image Segmentation
  • Monocular Depth Estimation (MiDaS/DPT)
  • Depth-Image-Based Rendering
  • OCR (Tesseract)
  • Image Preprocessing & Enhancement

Generative AI / LLMs

  • Diffusion & Image-Generation Workflows
  • Vision-Language Models
  • Prompt Engineering
  • RAG
  • LLM Evaluation
  • Hallucination Detection
  • Instruction & Preference Data Curation

Data & Deployment

  • NumPy
  • Pandas
  • Matplotlib
  • Seaborn
  • FastAPI
  • Flask
  • Streamlit
  • REST APIs
  • Docker
  • Git
  • Linux
  • MySQL
  • MongoDB
  • CUDA

03 — Track record

Experience

  1. AI Specialist — Computer Vision & Multimodal Evaluation (Contract)

    Aug 2026 — Present

    Snorkel AI · Remote

    • Evaluate Vision-Language Model outputs on image and video tasks, catching hallucinations, spatial-reasoning errors and wrong object/relationship claims, and write corrective annotations to improve the models.
    • Analyse detection, segmentation, depth, geometry and image-comparison tasks, breaking complex scenes into parts to pinpoint where a model fails, and build multimodal training/evaluation data for visual reasoning.
  2. Freelance AI/ML & AI Evaluation Specialist

    2024 — Present

    Contract work for AI training and evaluation platforms · Remote

    • Author expert-level evaluation tasks, rubrics and reference solutions for LLMs and AI agents, including image-based reasoning, mathematics and agentic tool-use scenarios.
    • Grade model responses against rubrics, flag reasoning and factual errors, and write preference/correction data for training.
  3. Lecturer — Computer Science

    Jul 2025 — Dec 2025

    RSR Rungta College of Engineering and Technology (RSR-RCET), Bhilai · On-site

    • Taught Image Processing, Cybersecurity, Cryptography and C to UG/PG students; mentored AI capstone projects.
    • Built an academic timetable system that automatically resolves faculty, classroom and period clashes for about 35 faculty, replacing manual spreadsheets.
  4. Junior Computer Vision Engineer

    Feb 2025 — Jun 2025

    CLAW Legaltech Private Limited · Remote

    • Built a generative-AI textile design system that creates custom fashion prints and patterns from user inputs, using PyTorch and image-generation models.
    • Added preprocessing, enhancement and style-variation controls, turning the research prototype into a product feature, and iterated with product stakeholders on quality and consistency.
  5. Junior Data Scientist

    Jan 2024 — Jun 2024

    Logiciel Analytics Private Limited · Remote

    • Built an end-to-end OCR pipeline (Python, Tesseract, OpenCV, Pandas) that turns scanned invoices into structured, queryable records.
    • Improved recognition accuracy with noise removal, adaptive thresholding, deskewing and contrast enhancement; automated a manual data-entry workflow.

04 — Selected work

Projects

2D-to-3D Video Conversion using Monocular Depth Estimation

M.Sc. Dissertation
  • Built a pipeline that converts standard 2D video into stereoscopic 3D using monocular depth estimation and depth-image-based rendering (DIBR).
  • Benchmarked MiDaS DPT-Large vs DPT-Hybrid (depth quality vs speed); synthesised left/right views with pixel warping and hole-filling, and used GPU acceleration and temporal smoothing to reduce flicker.
PythonPyTorchOpenCVNumPyCUDA

Generative Textile Design System

CLAW Legaltech
  • Generative-AI system that creates custom fashion prints and patterns from user inputs.
  • Preprocessing, enhancement and style-variation controls took the prototype to a shipped product feature.
PyTorchDiffusion ModelsOpenCVPython

OCR Invoice Extraction Pipeline

Logiciel Analytics
  • End-to-end pipeline turning scanned invoices into structured, queryable records.
  • Noise removal, adaptive thresholding, deskewing and contrast enhancement lifted recognition accuracy and removed a manual data-entry step.
TesseractOpenCVPandasPython

05 — Background

Education & certifications

2023 — 2025

M.Sc. Information Technology (AI & Machine Learning)

IIIT Lucknow

CGPA 8.00 / 10

2018 — 2021

B.Sc. Mathematics

Veer Kunwar Singh University

Certifications & achievements

06 — Say hello

Get in touch

Open to AI / Computer Vision roles and evaluation contracts — remote or with relocation. Based in Sasaram, Bihar.