I build intelligent systems at both ends of the stack — the algorithm that decides, and the silicon it decides on.
Dual-degree student building efficient intelligent systems — the kind that have to survive both a loss curve and a timing report.
Most people pick a side: the model or the machine. The interesting problems live in the seam.
🧠 The model — world models · model-based RL · GNNs · agentic systems · quantum ML | ⚙️ The machine — RTL · verification · AI accelerators · quantized & edge inference
Profile summary — click to expand
class Engineer:
"""The seam between what thinks and what it thinks on."""
name = "Tanush Pavan"
degrees = ["BS Data Science — IIT Madras",
"B.Tech Electrical & Electronics — Amrita Vishwa Vidyapeetham"]
thesis = "intelligence is an architecture problem before it is a scale problem"
experience = [("Wipro", "AI Engineer Intern", "applied ML, RL for sequential decisions"),
("ADRIN — ISRO", "Edge AI Intern", "HLS datapath for MobileViTv2 on Versal VCK-190")]
research = {"world_models": ["RSSM", "latent imagination", "learned dynamics"],
"rl": ["model-based", "actor-critic", "λ-returns"],
"graphs": ["message passing", "GNNs over netlists"],
"agentic": ["planner/critic loops", "tool use", "cost-aware search"],
"quantum": ["variational circuits", "parameter-shift", "QML"],
"verification": ["UVM", "cocotb", "SVA", "formal model checking"],
"silicon": ["RTL", "microarchitecture", "accelerators", "HLS"]}
def philosophy(self):
return ("Efficiency is not an optimization pass you run at the end. "
"It is a decision you make in the first hour, and defend in every one after.")
This panel is not written by me. Every night a GitHub Action reads what I actually pushed,
compresses it into a state vector, and performs one imagination rollout from that latent —
then commits the result. A world model that dreams about the person building it.
🌙 How the dream works - The Nightly Dreamer model
The loop. A scheduled workflow wakes tools/dream.py at 02:30 IST. The script pulls my recent public GitHub events, classifies commits and repo names into research threads, and measures cadence — how many pushes, which threads dominate, what hour of the night the commits landed. That becomes a small state vector, shown as chips along the bottom of the panel so you can see the evidence the dream was conditioned on. A model then performs one rollout from that state, and the result is typeset into the SVG above and committed.
Why nightly and committed, rather than generated per visitor. The obvious design is a serverless endpoint that runs inference on every page view. It doesn't work: GitHub proxies every README image through Camo and caches it hard, so the endpoint serves one request and Camo serves everyone else a stale copy. You'd pay per cache-miss, expose a public inference endpoint, and risk Camo timing out mid-render. A committed artefact loads instantly, costs about thirty calls a month, and keeps the key in GitHub Secrets.
The log is the point. Every rollout is a commit, so dreams/log.json accumulates an archaeological record — what the repo was thinking in March, which threads were hot, which nights ran long. That history is more interesting than an ephemeral greeting, and it's reviewable: anything the model says can be read, reverted, or argued with.
It degrades, it doesn't break. With no API key the script composes the rollout deterministically from the same state vector. The panel always renders; only the prose gets less surprising.
|
Role: Edge AI / HW-SW co-design
Work:
- Deployed an INT8 MobileViTv2-XS vision
transformer on AMD Versal VCK190
(Vitis AI / HLS / Vivado)
- Sub-100 ms onboard Earth-observation
inference for satellite systems
- Latency / throughput / utilization
trade-offs across Versal and Zynq
Takeaway: >
A transformer on edge silicon is a memory
hierarchy problem wearing an attention mask. |
Role: Applied AI / ML engineering
Work:
- Built a production spam/ham classifier
(TF-IDF → Random Forest + LR meta-model)
for high-volume enterprise streams
- Re-architected the real-time inference
pipeline — batching + vectorized preproc
Takeaway: >
Cut processing latency 13.4% at scale —
the model was never the slow part. |
|
Designed and verified a patent-pending 64×64 systolic-array binary neural network accelerator in SystemVerilog for energy-efficient edge inference, with full RTL verification and performance analysis. |
More recent builds
| Build | What it is | Result |
|---|---|---|
| Agentic RISC-V Verification (AVA · VOE) | 12-agent pipeline: RTL → testbench → golden-ISS diff → coverage → new tests, with a frozen kernel where a claim needs a re-checkable witness | the green result that means nothing, caught |
| Anvil | Hardware-aware LLM quantization agent (Arm AI Challenge) — surrogate-guided search over 2.7×10⁸ per-layer configs | greedy's 60-eval result in ~9 · 1.8× faster TTFT |
| GridSetu | City-grid simulator (day-ahead market, AC power flow) turning a community battery into a forecast-backed reliability reserve | −80% evening unserved energy in simulation |
| DriftSense | SEM-image patch localization via FFT + line-edge-roughness fingerprinting (no neural nets) | 3.3× better than classical NCC · MunichTech winner |
| GeoReach | Sentinel-1 SAR + OpenStreetMap flood-accessibility mapping for Assam | IEEE GRSS hackathon winner |
| Competition | Result | Domain |
|---|---|---|
| IEEE GRSS · SAADRI | 🥇 Winner | GeoReach — Sentinel-1 SAR · OpenStreetMap · flood mapping |
| MunichTech EXPO 2026 | 🥇 Winner | DriftSense — OpenCV · SEM imaging · physics-based localization |
| Smart India Hackathon 2025 | 🏅 Grand Finale — national finalist | Govt. of India, national-scale problem statement |
| Analog Circuit Design Challenge | 🏅 Finalist | IIT Madras — analog / mixed-signal design |
| Mirabilis Design Hacks | 🏅 Finalist | System-level modelling & architecture |
Two wins and three national finals across geospatial AI, semiconductor CV, analog design and system architecture.
The through-line isn't the subject. It's going from a cold problem statement to a defensible build in 36 hours.
Developer-to-developer write-ups — case studies, project teardowns, technical deep-dives and data stories.
The reasoning behind the work, not just the result.
|
🔬 Case Studies Real problems, the approach, and the number that mattered |
🧩 Project Write-Ups How a repo was built and what I'd change |
⚙️ Technical Blogs Focused deep-dives on one idea |
📊 Data Stories A dataset, a question, the answer |
Latest — MobileViTv2 on the Versal VCK-190 · Why hold violations survive simulation
Everything above is the what. Below is the how — the research, the maths, and the silicon,
each folded into a collapsible. Open what interests you; skip what doesn't.
Six threads, one question: how do you build something that models its world well enough to act in it — cheaply enough to matter?
The formalism I actually work in. Not decoration — these are the objects I'm debugging when something doesn't converge.
🌍 World Models & Model-Based RL — a latent you can plan inside
A world model compresses observations into a latent whose dynamics are learnable. The RSSM splits that latent in two: a deterministic path
Training maximises the ELBO shown in the panel above — reconstruct the observation, predict the reward, and pay a KL price for every bit of surprise smuggled into the latent. Once the dynamics are learned you never touch the environment again: you roll out inside the model and train on imagined trajectories, bootstrapped with a
The part that's genuinely hard:
Why model-based at all: model-free RL pays for every gradient step in real environment interactions. A world model converts sample complexity into compute complexity — and compute is the thing I know how to make cheaper in hardware. That sentence is the whole reason this profile has both halves.
🕸️ Graph Neural Networks — because a netlist is a graph
Message passing in its general form — aggregate from the neighbourhood, update, repeat, with
Swap that aggregator for attention and you get a GAT; normalise it spectrally and you get a GCN. They are the same idea wearing different clothes, which is also why transformers keep showing up in this literature.
Why I care: a gate-level netlist, a placement, a routing congestion map and a dataflow graph are all graphs. Every EDA problem that currently costs hours of heuristic search is a graph learning problem nobody has finished attacking.
🤖 Agentic AI — planning with a budget
An agent that calls tools is doing sequential decision-making where actions have cost, not just consequence. The honest objective includes the bill:
where
Where I've shipped this: the agentic RISC-V verification platform — twelve agents, most running with no EDA toolchain, where every tool call has a real cost. And Anvil, a quantization agent whose surrogate-guided planner matched a 60-evaluation greedy search in ~9 on-device evals. Both make you take
⚡ Quantized & Binary Networks — the math that makes edge inference possible
Uniform affine quantization maps a float to an integer grid through a scale
At the binary extreme, weights collapse to a sign and one scale per filter:
which turns a multiply-accumulate into XNOR + popcount — and that is a sentence about hardware, not about machine learning.
The constraint that decides everything — arithmetic intensity against the roofline:
Most "slow" models are not compute-bound. They are memory-bound, and quantization is a bandwidth optimization that happens to look like a numerics one.
⚛️ The quantum thread — what I'm actually studying
A state on
Gradients come from the parameter-shift rule — exact, not finite-difference, which is the detail that makes it trainable at all:
The honest caveat: barren plateaus. For a random deep ansatz on
Why it sits next to the silicon work: both are about extracting useful computation from a physical substrate that does not care about your abstractions. Coherence times and setup time are the same genre of constraint.
⏱️ Timing, power, area — the three numbers you are always trading
Setup closure — the clock period has to cover the whole combinational path:
Hold closure — the failure mode that survives simulation and kills silicon, because it is frequency-independent:
You cannot slow the clock to fix a hold violation. That asymmetry is why hold buffers exist and why CTS matters more than it looks.
Power splits into switching,
Amdahl, for accelerator scoping — the sentence that kills bad accelerator proposals early:
If your kernel is 60% of runtime, an infinitely fast accelerator buys you 2.5×. Profile before you build.
🔧 What I actually do at each stage
| Stage | What I do | Tooling |
|---|---|---|
| Spec / architecture | throughput, latency and area budgets before a line of RTL; roofline analysis on the target kernel | Python, spreadsheets, arguing |
| RTL | synthesisable SystemVerilog, clean handshakes, parameterised datapaths | SystemVerilog, Verilog |
| Lint / CDC | rule cleanliness and clock-domain safety before simulation gets expensive | lint flows, CDC review |
| Verification | UVM environments, cocotb testbenches, SVA properties, coverage closure | UVM, cocotb, SVA |
| Synthesis | constraint writing, timing exploration, understanding what the tool did to my intent | Vivado, standard flows |
| FPGA / HLS | HLS datapath design and hardware-aware model restructuring | Vivado HLS, Versal ACAP |
| Signoff literacy | reading STA reports, understanding DRC/LVS as a design constraint rather than someone else's problem | reports, and patience |
Concretely: at ADRIN (ISRO) I designed a custom HLS-based datapath to deploy MobileViTv2 — a vision transformer — on the AMD Versal ACAP VCK-190. Transformers on edge silicon are an exercise in memory hierarchy, not in FLOPs. The attention block is the bandwidth problem; everything else is arithmetic you can schedule.
🔍 Why constrained-random needs so many runs — and when to stop
Random stimulus hitting
For
The complementary argument, for formal: model checking does not sample the state space, it quantifies over it. Where a property is small and the state space is bounded, a proof beats
My working rule: lint and CDC first because they're free. cocotb for iteration speed while the design is still moving. UVM once the interfaces stabilise and you need real reuse. SVA everywhere, because an assertion that fires next to the bug is worth a hundred waveforms. Formal on the control logic where the state space is small enough to be exhausted.
🐍 cocotb — and the open-source tooling around it
Python testbenches are not a toy. They give you the whole scientific stack next to your DUT: numpy for reference models, pytest for structure, CI that actually runs on every push. The tradeoff is simulation speed at the boundary — which matters less than people assume for anything below full-chip.
I maintain cocotb-v2-migration-helper for exactly this reason: the v1 → v2 API break is mechanical enough to automate and tedious enough that people put it off, and testbench debt compounds faster than design debt.
🔬 Process physics — the constraints that reach all the way up to RTL
Lithography sets the floor. The panel above carries the Rayleigh criterion: EUV at
Yield decides whether any of it matters. Murphy's model, for a die of area
Yield falls off superlinearly with die area. That one fact is why chiplets exist, why big dies are disproportionately expensive, and why "just make the accelerator bigger" is an economic proposal before it is an architectural one.
Devices, as they've actually moved: planar → FinFET → gate-all-around nanosheet, each transition a geometric answer to the gate losing control of a shrinking channel.
Why an RTL person should know this: wire delay does not scale like gate delay. Past a certain node interconnect dominates — so floorplan is a microarchitectural decision, not a backend one, and locality in your dataflow is worth more than gate count.
Original work in progress. Written down here because a claim you've committed to a public repo
is a claim you have to keep honest.
| Thread | Hypothesis | Status |
|---|---|---|
| Hardware-aware latent dynamics | world-model latents shaped by a hardware cost term learn representations that are cheaper to run, not just cheaper to store | active |
| GNNs over netlists | structural graph learning can replace heuristic passes in verification triage and design-space search | active |
| Cost-aware agentic planning | making |
active, from the verification-agent work |
| Binary/quantized accelerator co-design | quantization schemes chosen jointly with the datapath beat schemes chosen for the model alone | active, tied to ADRIN work |
| Verification-informed architecture | designs that are cheap to verify are a distinguishable class, and the property is predictable from the RTL | early |
| Credential | Issuer | Why it's here |
|---|---|---|
| Hardware Security | University of Maryland | trojans, side channels, and PUFs — the attacks that live below the software threat model |
| VLSI Design | L&T EduTech | industry-framed digital design flow, end to end |
| Algorithms Specialization | Stanford University | the four-course track — divide & conquer, graphs, greedy/DP, NP-completeness |
Agentic AI Foundations Associate 1Z0-1157-26 |
Oracle | agent reasoning patterns, tool orchestration, MCP |
| AI Agents Course | Hugging Face | building and evaluating agents against real benchmarks |
SQL AI Developer Associate DP-800 |
Microsoft | vector search, embeddings and RAG pushed down into the database engine |
Certified Data Engineer – Associate DEA-C01 |
Amazon Web Services | pipelines, storage and orchestration — the layer the models actually run on |
GitHub Foundations GH-900 |
GitHub | the collaboration plumbing this whole profile is built on |
| Google Data Analytics Professional | the eight-course analytics track |
Hardware Security and the Algorithms specialization are the two that show up most in my actual work —
one because attacks find the layer you forgot to model, the other because complexity analysis is
the only honest way to argue about a design before you've built it.
BS in Data Science — Indian Institute of Technology Madras
B.Tech in Electrical & Electronics Engineering — Amrita Vishwa Vidyapeetham
Two degrees, deliberately. One taught me to reason about data and uncertainty;
the other taught me what actually happens when electrons have to carry the answer.
Everything above this line is hand-authored SVG. No template, no generator, no screenshot.
The page is the portfolio piece — so here is the design system it runs on.
🎨 The design system, and how the page is built
Colour tokens — one ramp from substrate to signal, with a violet axis reserved for anything quantum or probabilistic. substrate #030811 → die #071226 → well #0C1D3E → trace #1B3566, accented by signal #38BDF8 and signal-alt #22D3EE, text in ink #C8E0F8 and ink-dim #7FA6D4, probabilistic domains in quantum #A78BFA / #C084FC.
Type scale — three families, three jobs. Segoe UI for headings, JetBrains Mono for anything machine-adjacent, a math serif for equations. Section headers at 30px / 700 / +10 letter-spacing; annotations at 9–11px mono with lowered opacity, so they read as marginalia rather than content.
Motion, deliberately restrained — every animation is a 3–9 second loop with no easing spikes, so nothing competes for attention. Travelling pulses along circuit rails, a slow scan across banners, coverage bars that fill once and freeze, a Bloch vector that precesses. Motion signals aliveness, not urgency.
Layout — a fixed 1400px design width, 16–18px radii, 60px gutters, and a 10%-opacity dot grid on every panel so the page reads as one continuous surface rather than a stack of unrelated images.
Accessibility — every graphic carries alt text, nothing critical is colour-only, body contrast clears 7:1, and the page degrades to readable structured markdown if images fail.
Built from source, not by hand — every banner, divider and diagram is generated by parameterised Python in tools/, so the geometry is computed rather than eyeballed. Equations are typeset to SVG with MathJax and served light/dark through <picture>, because GitHub mangles double-dollar math inside collapsible blocks — it strips backslashes and eats underscore pairs as italics. The TeX source lives in tools/equations.json and in every image's alt. GitHub Actions regenerate the contribution graph nightly, and the dream with it.
Got a question about any of this? World models, RTL, verification, the edge-AI work, or how this page is built — ask in the open.
Answered in the open, so the next person with the same question can read it too.
A live, interactive Q&A widget is coming to the portfolio site.
Open to conversations about
world models and model-based RL · AI accelerator architecture · verification methodology
graph learning for EDA · quantized and binary inference · anything at the hardware/ML seam
Building efficient, intelligent systems where hardware, software, and learning meet.

