Building machine intelligence with a clear signal.

I design and build practical AI systems — from machine learning experiments and RAG pipelines to data-driven products that turn messy information into usable decisions.

About

Builder, tinkerer, AI engineer.

Who I am

I'm Sanjay Ram A, an AI engineering student from Tamil Nadu building real products on the side. I started with curiosity about how models actually work — and ended up shipping apps, fine-tuning LLMs, and founding Dotwellabs.

My work sits at the intersection of applied machine learning, mobile AI, and product thinking. I care about systems that are measurable, reproducible, and useful outside a notebook — not just impressive in a demo.

When I'm not training models or debugging inference pipelines, I'm writing about what I learn, contributing to open source, and looking for problems worth solving with AI.

4+ Projects shipped — from Play Store apps to hackathon winners
2+ Years building AI systems, models, and data products
Models and datasets published on Hugging Face
Open Available for AI engineering roles, research collabs, and ML projects

Selected Work

Projects with working models, not decorative buzzwords.

P/001
LLMs MCP Agentic Skills Agent Loops

AnyLLM

AnyLLM is the ultimate AI chat app — one unified workspace to talk to GPT-4o, Claude, Gemini, Llama, Groq, and more. Stop juggling five different apps. Bring your own API key and get the most out of every top AI model in one place. Supports Web search, MCPs, AI-Debates, Custom providers via OpenAI-compatible APIs and Agentic skills.

P/002
Fine-tuning PyTorch Flutter Catus Edge-AI

Fixgemma: AI Repair Assistant

AI-powered appliance repair assistant, built entirely on-device for the Gemma 4 Good Hackathon. FixGemma is a Flutter application that leverages Gemma 4 (e2b & e4b models) via the Cactus AI inference engine to provide users with private, completely offline, multimodal repair assistance.

P/003
RAG Ollama LLMs

AI Debate Simulator

An interactive platform that enables AI agents from different domains to engage in structured debates on various topics. it is a Flask-based web application that facilitates debates between AI agents specialized in different domains. Each agent utilizes domain-specific knowledge bases and Large Language Models (LLMs) to generate informed responses and engage in meaningful discussions.

P/004
ONNX Voice Cloning Edge-AI Flutter

Pocket Speech: Voice Cloning TTS

Offline mobile TTS pipeline built around Kyutai’s Pocket TTS model via ONNX runtime (sherpa-onnx), with zero-shot voice cloning fully on-device. Record a 5–12s Voice Profile or use built-in voices, tune inference parameters for speed vs quality — fast synthesis with no network dependency and complete privacy.

P/005
PyTorch Llama Architecture LLM Pretraining

NanoAgent-15M

Pretrained a 15M-parameter Llama-style decoder-only transformer from scratch on a single RTX 3050 (4GB VRAM) across ∼1B tokens, with a custom 6K byte-level BPE tokenizer and staged SFT for chat and tool-calling. Final validation perplexity ∼53.4 at under 1.1GB peak memory; open-sourced model and full pipeline.

Technical Stack

Tools I use to move from notebook to product.

01

Models & Training

  • PyTorch
  • Transformers
  • Fine-tuning — SFT / PEFT
  • Unsloth
  • Diffusion / ViT

02

Retrieval & Agents

  • RAG / Vector Search
  • LangChain / LlamaIndex
  • AI Agents / MCP
  • Prompt Engineering
  • Embeddings

03

Edge & Mobile

  • ONNX / Quantization
  • On-Device Deployment
  • Cactus / CoreML / MediaPipe
  • Flutter / Android
  • Hugging Face

04

Backend & Tools

  • Python / SQL
  • FastAPI / Flask
  • Docker / Git / Linux
  • MLflow / Streamlit / Gradio
  • HTML / CSS / JS

Interests

The areas I keep returning to.

01

Applied LLM Systems

Finetuning, RAG, agents, evaluation, prompt design, and reliable AI workflows.

02

LLM Training

Dataset curation, supervised fine-tuning (SFT), parameter-efficient tuning, and rigorous evaluation.

03

Data Products

Dashboards, APIs, and interfaces that make model outputs understandable.

04

Model Evaluation

Metrics, error analysis, reproducibility, and honest measurement.

Blog / Notes

Short notes from experiments, failures, and rebuilds.

FixGemma: I Squeezed Gemma 4 Onto an $80 Phone to Fight E-Waste

Kaggle Gemma hackathon writeup — fine-tuning E4B/E2B, running fully offline on a budget phone, and shipping a carousel UI.

Read note
Sanjay Ram A, AI Engineer — portrait photo
Sanjay Ram A — AI Engineer

Let’s build something that learns clearly.

I’m interested in AI engineering roles, ML projects, research collaborations, and product ideas where machine learning needs to be useful, measured, and shipped.