Skip to content
View HUNT-001's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report HUNT-001

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
HUNT-001/README.md

Tanush Pavan

Typing SVG

GitHub   LinkedIn   Portfolio   Email

I build intelligent systems at both ends of the stack — the algorithm that decides, and the silicon it decides on.

Profile Views Followers BNN accelerator Hackathon wins SIH 2025

Boot Sequence

About Me

The model, the machine, and the seam between them

Dual-degree student building efficient intelligent systems — the kind that have to survive both a loss curve and a timing report.
Most people pick a side: the model or the machine. The interesting problems live in the seam.

🧠 The model — world models · model-based RL · GNNs · agentic systems · quantum ML   |   ⚙️ The machine — RTL · verification · AI accelerators · quantized & edge inference

 Profile summary — click to expand
class Engineer:
    """The seam between what thinks and what it thinks on."""

    name    = "Tanush Pavan"
    degrees = ["BS Data Science — IIT Madras",
               "B.Tech Electrical & Electronics — Amrita Vishwa Vidyapeetham"]
    thesis  = "intelligence is an architecture problem before it is a scale problem"

    experience = [("Wipro",        "AI Engineer Intern", "applied ML, RL for sequential decisions"),
                  ("ADRIN — ISRO", "Edge AI Intern",     "HLS datapath for MobileViTv2 on Versal VCK-190")]

    research = {"world_models": ["RSSM", "latent imagination", "learned dynamics"],
                "rl":           ["model-based", "actor-critic", "λ-returns"],
                "graphs":       ["message passing", "GNNs over netlists"],
                "agentic":      ["planner/critic loops", "tool use", "cost-aware search"],
                "quantum":      ["variational circuits", "parameter-shift", "QML"],
                "verification": ["UVM", "cocotb", "SVA", "formal model checking"],
                "silicon":      ["RTL", "microarchitecture", "accelerators", "HLS"]}

    def philosophy(self):
        return ("Efficiency is not an optimization pass you run at the end. "
                "It is a decision you make in the first hour, and defend in every one after.")

Mission & Domains

Signal Log

Dream Log

The most recent latent rollout

This panel is not written by me. Every night a GitHub Action reads what I actually pushed,
compresses it into a state vector, and performs one imagination rollout from that latent —
then commits the result. A world model that dreams about the person building it.

 🌙  How the dream works - The Nightly Dreamer model

The loop. A scheduled workflow wakes tools/dream.py at 02:30 IST. The script pulls my recent public GitHub events, classifies commits and repo names into research threads, and measures cadence — how many pushes, which threads dominate, what hour of the night the commits landed. That becomes a small state vector, shown as chips along the bottom of the panel so you can see the evidence the dream was conditioned on. A model then performs one rollout from that state, and the result is typeset into the SVG above and committed.

Why nightly and committed, rather than generated per visitor. The obvious design is a serverless endpoint that runs inference on every page view. It doesn't work: GitHub proxies every README image through Camo and caches it hard, so the endpoint serves one request and Camo serves everyone else a stale copy. You'd pay per cache-miss, expose a public inference endpoint, and risk Camo timing out mid-render. A committed artefact loads instantly, costs about thirty calls a month, and keeps the key in GitHub Secrets.

The log is the point. Every rollout is a commit, so dreams/log.json accumulates an archaeological record — what the repo was thinking in March, which threads were hot, which nights ran long. That history is more interesting than an ephemeral greeting, and it's reviewable: anything the model says can be read, reverted, or argued with.

It degrades, it doesn't break. With no API key the script composes the rollout deterministically from the same state vector. The panel always renders; only the prose gets less surprising.

Experience

ADRIN — NRSC, ISRO

Edge AI Research Intern · 2026

Role: Edge AI / HW-SW co-design
Work:
  - Deployed an INT8 MobileViTv2-XS vision
    transformer on AMD Versal VCK190
    (Vitis AI / HLS / Vivado)
  - Sub-100 ms onboard Earth-observation
    inference for satellite systems
  - Latency / throughput / utilization
    trade-offs across Versal and Zynq
Takeaway: >
  A transformer on edge silicon is a memory
  hierarchy problem wearing an attention mask.

WIPRO

AI-ML Engineering Intern · 2025

Role: Applied AI / ML engineering
Work:
  - Built a production spam/ham classifier
    (TF-IDF → Random Forest + LR meta-model)
    for high-volume enterprise streams
  - Re-architected the real-time inference
    pipeline — batching + vectorized preproc
Takeaway: >
  Cut processing latency 13.4% at scale —
  the model was never the slow part.

🔬 Research · BNN Systolic Accelerator

Hardware Research Engineer · 2025 – 2026 · Coimbatore

Designed and verified a patent-pending 64×64 systolic-array binary neural network accelerator in SystemVerilog for energy-efficient edge inference, with full RTL verification and performance analysis.

Projects

More recent builds

Build What it is Result
Agentic RISC-V Verification (AVA · VOE) 12-agent pipeline: RTL → testbench → golden-ISS diff → coverage → new tests, with a frozen kernel where a claim needs a re-checkable witness the green result that means nothing, caught
Anvil Hardware-aware LLM quantization agent (Arm AI Challenge) — surrogate-guided search over 2.7×10⁸ per-layer configs greedy's 60-eval result in ~9 · 1.8× faster TTFT
GridSetu City-grid simulator (day-ahead market, AC power flow) turning a community battery into a forecast-backed reliability reserve −80% evening unserved energy in simulation
DriftSense SEM-image patch localization via FFT + line-edge-roughness fingerprinting (no neural nets) 3.3× better than classical NCC · MunichTech winner
GeoReach Sentinel-1 SAR + OpenStreetMap flood-accessibility mapping for Assam IEEE GRSS hackathon winner

Achievements

Competition record

Competition Result Domain
IEEE GRSS · SAADRI 🥇 Winner GeoReach — Sentinel-1 SAR · OpenStreetMap · flood mapping
MunichTech EXPO 2026 🥇 Winner DriftSense — OpenCV · SEM imaging · physics-based localization
Smart India Hackathon 2025 🏅 Grand Finale — national finalist Govt. of India, national-scale problem statement
Analog Circuit Design Challenge 🏅 Finalist IIT Madras — analog / mixed-signal design
Mirabilis Design Hacks 🏅 Finalist System-level modelling & architecture

Two wins and three national finals across geospatial AI, semiconductor CV, analog design and system architecture.
The through-line isn't the subject. It's going from a cold problem statement to a defensible build in 36 hours.

Writing

Developer-to-developer write-ups — case studies, project teardowns, technical deep-dives and data stories.
The reasoning behind the work, not just the result.

🔬 Case Studies

Real problems, the approach, and the number that mattered

🧩 Project Write-Ups

How a repo was built and what I'd change

⚙️ Technical Blogs

Focused deep-dives on one idea

📊 Data Stories

A dataset, a question, the answer

Write-ups   Portfolio blog

Latest — MobileViTv2 on the Versal VCK-190  ·  Why hold violations survive simulation

Technical Depth

Everything above is the what. Below is the how — the research, the maths, and the silicon,
each folded into a collapsible. Open what interests you; skip what doesn't.

Research

Research Surface

Six threads, one question: how do you build something that models its world well enough to act in it — cheaply enough to matter?

Mathematics

RSSM Variational Objective

The formalism I actually work in. Not decoration — these are the objects I'm debugging when something doesn't converge.

 🌍  World Models & Model-Based RL — a latent you can plan inside

A world model compresses observations into a latent whose dynamics are learnable. The RSSM splits that latent in two: a deterministic path $h_t$ carrying memory, and a stochastic path $z_t$ carrying uncertainty.

h_t = f_\theta\!\left(h_{t-1},\, z_{t-1},\, a_{t-1}\right), \qquad z_t \sim q_\phi\!\left(z_t \mid h_t,\, o_t\right)

Training maximises the ELBO shown in the panel above — reconstruct the observation, predict the reward, and pay a KL price for every bit of surprise smuggled into the latent. Once the dynamics are learned you never touch the environment again: you roll out inside the model and train on imagined trajectories, bootstrapped with a $\lambda$-return.

V^{\lambda}_t \;=\; r_t \;+\; \gamma\Big[(1-\lambda)\, v_\psi(s_{t+1}) \;+\; \lambda\, V^{\lambda}_{t+1}\Big]

The part that's genuinely hard: $\beta$ on that KL term is the whole argument. Too low and the model memorises pixels instead of learning consequence. Too high and the posterior collapses onto the prior — the model stops dreaming, and every rollout returns the same beige future.

Why model-based at all: model-free RL pays for every gradient step in real environment interactions. A world model converts sample complexity into compute complexity — and compute is the thing I know how to make cheaper in hardware. That sentence is the whole reason this profile has both halves.

 🕸️  Graph Neural Networks — because a netlist is a graph

Message passing in its general form — aggregate from the neighbourhood, update, repeat, with $\bigoplus$ any permutation-invariant aggregator:

h_v^{(k+1)} \;=\; \sigma\!\left( W^{(k)} h_v^{(k)} \;+\; \bigoplus_{u \in \mathcal{N}(v)} \frac{1}{c_{vu}}\, M^{(k)}\!\left(h_u^{(k)},\, e_{uv}\right) \right)

Swap that aggregator for attention and you get a GAT; normalise it spectrally and you get a GCN. They are the same idea wearing different clothes, which is also why transformers keep showing up in this literature.

Why I care: a gate-level netlist, a placement, a routing congestion map and a dataflow graph are all graphs. Every EDA problem that currently costs hours of heuristic search is a graph learning problem nobody has finished attacking.

 🤖  Agentic AI — planning with a budget

An agent that calls tools is doing sequential decision-making where actions have cost, not just consequence. The honest objective includes the bill:

a^{*} \;=\; \arg\max_{a \in \mathcal{A}} \; \mathbb{E}_{s' \sim T}\!\left[\, R(s, a, s') + \gamma V(s') \,\right] \;-\; \lambda\, c(a)

where $c(a)$ is latency, tokens, API spend, or blast radius. Drop that term and you get an agent that solves the task by brute-force calling everything it can reach. Add tree search with a learned value on top and you are doing planning, not prompting.

Where I've shipped this: the agentic RISC-V verification platform — twelve agents, most running with no EDA toolchain, where every tool call has a real cost. And Anvil, a quantization agent whose surrogate-guided planner matched a 60-evaluation greedy search in ~9 on-device evals. Both make you take $c(a)$ seriously.

 ⚡  Quantized & Binary Networks — the math that makes edge inference possible

Uniform affine quantization maps a float to an integer grid through a scale $s$ and zero-point $z$, then clips. Rounding has zero gradient almost everywhere, so training leans on the straight-through estimator — pretend the quantizer is the identity inside the clipping range:

\frac{\partial \hat{x}}{\partial x} \;\approx\; \mathbf{1}_{\left\{\, q_{\min} \,\le\, x/s + z \,\le\, q_{\max} \,\right\}}

At the binary extreme, weights collapse to a sign and one scale per filter:

w_b = \mathrm{sign}(w), \qquad \alpha = \frac{\lVert W \rVert_1}{n}, \qquad W \approx \alpha\, w_b

which turns a multiply-accumulate into XNOR + popcount — and that is a sentence about hardware, not about machine learning.

The constraint that decides everything — arithmetic intensity against the roofline:

P_{\text{attainable}} = \min\!\left(P_{\text{peak}},\; I \cdot B_{\text{mem}}\right), \qquad I = \frac{\text{FLOPs}}{\text{Bytes moved}}

Most "slow" models are not compute-bound. They are memory-bound, and quantization is a bandwidth optimization that happens to look like a numerics one.

Quantum

Variational Quantum Eigensolver

 ⚛️  The quantum thread — what I'm actually studying

A state on $n$ qubits lives in $\mathbb{C}^{2^n}$, which is the entire promise and the entire problem. Variational algorithms hedge: a shallow parameterised circuit on the device, a classical optimiser outside it. The panel above is the whole method — the variational principle guarantees $E(\boldsymbol\theta)$ never undershoots the ground state, and the Hamiltonian decomposes into Pauli strings you can actually measure.

Gradients come from the parameter-shift rule — exact, not finite-difference, which is the detail that makes it trainable at all:

\frac{\partial E}{\partial \theta_i} \;=\; \frac{1}{2}\left[ E\!\left(\theta_i + \tfrac{\pi}{2}\right) - E\!\left(\theta_i - \tfrac{\pi}{2}\right) \right]

The honest caveat: barren plateaus. For a random deep ansatz on $n$ qubits, gradient variance decays as $\mathcal{O}(2^{-n})$ — the landscape flattens exponentially and the optimiser has nothing to descend. Structured ansätze and local cost functions are the live area, and I'm reading rather than claiming here.

Why it sits next to the silicon work: both are about extracting useful computation from a physical substrate that does not care about your abstractions. Coherence times and setup time are the same genre of constraint.

Silicon

RTL to GDSII flow

 ⏱️  Timing, power, area — the three numbers you are always trading

Setup closure — the clock period has to cover the whole combinational path:

T_{\text{clk}} \;\ge\; t_{cq} \;+\; t_{\text{logic,max}} \;+\; t_{\text{setup}} \;-\; t_{\text{skew}} \;+\; t_{\text{jitter}}

Hold closure — the failure mode that survives simulation and kills silicon, because it is frequency-independent:

t_{cq} \;+\; t_{\text{logic,min}} \;\ge\; t_{\text{hold}} \;+\; t_{\text{skew}}

You cannot slow the clock to fix a hold violation. That asymmetry is why hold buffers exist and why CTS matters more than it looks.

Power splits into switching, $\alpha C_L V_{DD}^2 f$, and leakage, $V_{DD} I_\text{leak}$. The quadratic on $V_{DD}$ is the single most exploitable fact in low-power design — and the reason DVFS beats almost any microarchitectural trick you can name.

Amdahl, for accelerator scoping — the sentence that kills bad accelerator proposals early:

S = \frac{1}{(1 - p) + \dfrac{p}{s}}

If your kernel is 60% of runtime, an infinitely fast accelerator buys you 2.5×. Profile before you build.

 🔧  What I actually do at each stage
Stage What I do Tooling
Spec / architecture throughput, latency and area budgets before a line of RTL; roofline analysis on the target kernel Python, spreadsheets, arguing
RTL synthesisable SystemVerilog, clean handshakes, parameterised datapaths SystemVerilog, Verilog
Lint / CDC rule cleanliness and clock-domain safety before simulation gets expensive lint flows, CDC review
Verification UVM environments, cocotb testbenches, SVA properties, coverage closure UVM, cocotb, SVA
Synthesis constraint writing, timing exploration, understanding what the tool did to my intent Vivado, standard flows
FPGA / HLS HLS datapath design and hardware-aware model restructuring Vivado HLS, Versal ACAP
Signoff literacy reading STA reports, understanding DRC/LVS as a design constraint rather than someone else's problem reports, and patience

Concretely: at ADRIN (ISRO) I designed a custom HLS-based datapath to deploy MobileViTv2 — a vision transformer — on the AMD Versal ACAP VCK-190. Transformers on edge silicon are an exercise in memory hierarchy, not in FLOPs. The attention block is the bandwidth problem; everything else is arithmetic you can schedule.

Verification

Verification methodology

 🔍  Why constrained-random needs so many runs — and when to stop

Random stimulus hitting $n$ distinct coverage bins is the coupon-collector problem. The expected number of runs to hit all of them:

\mathbb{E}[N] \;=\; n \sum_{k=1}^{n} \frac{1}{k} \;\approx\; n \ln n + \gamma n, \qquad \gamma \approx 0.5772

For $n = 1000$ bins that is roughly 7,500 runs — and the tail is worse than the mean suggests. This is the quantitative argument for directed tests on the hard corners: you do not random-walk into a 1-in-$10^6$ state, you go there deliberately.

The complementary argument, for formal: model checking does not sample the state space, it quantifies over it. Where a property is small and the state space is bounded, a proof beats $7{,}500$ simulations and finishes sooner.

My working rule: lint and CDC first because they're free. cocotb for iteration speed while the design is still moving. UVM once the interfaces stabilise and you need real reuse. SVA everywhere, because an assertion that fires next to the bug is worth a hundred waveforms. Formal on the control logic where the state space is small enough to be exhausted.

 🐍  cocotb — and the open-source tooling around it

Python testbenches are not a toy. They give you the whole scientific stack next to your DUT: numpy for reference models, pytest for structure, CI that actually runs on every push. The tradeoff is simulation speed at the boundary — which matters less than people assume for anything below full-chip.

I maintain cocotb-v2-migration-helper for exactly this reason: the v1 → v2 API break is mechanical enough to automate and tedious enough that people put it off, and testbench debt compounds faster than design debt.

Fabrication

Fabrication cross-section

 🔬  Process physics — the constraints that reach all the way up to RTL

Lithography sets the floor. The panel above carries the Rayleigh criterion: EUV at $\lambda = 13.5\,\text{nm}$ with $\mathrm{NA} = 0.33$ reaches roughly 13 nm half-pitch in one exposure. Below that you multi-pattern — LELE, SADP, SAQP, each adding cost, mask count and overlay error — or move to High-NA and accept a smaller field. Note the square on NA in the depth-of-focus term: every gain in resolution costs focus budget quadratically. There is no free tightening.

Yield decides whether any of it matters. Murphy's model, for a die of area $A$ and defect density $D_0$:

Y = \left( \frac{1 - e^{-A D_0}}{A D_0} \right)^{2}

Yield falls off superlinearly with die area. That one fact is why chiplets exist, why big dies are disproportionately expensive, and why "just make the accelerator bigger" is an economic proposal before it is an architectural one.

Devices, as they've actually moved: planar → FinFET → gate-all-around nanosheet, each transition a geometric answer to the gate losing control of a shrinking channel.

Why an RTL person should know this: wire delay does not scale like gate delay. Past a certain node interconnect dominates — so floorplan is a microarchitectural decision, not a backend one, and locality in your dataflow is worth more than gate count.

Algorithms

Original work in progress. Written down here because a claim you've committed to a public repo
is a claim you have to keep honest.

Thread Hypothesis Status
Hardware-aware latent dynamics world-model latents shaped by a hardware cost term learn representations that are cheaper to run, not just cheaper to store active
GNNs over netlists structural graph learning can replace heuristic passes in verification triage and design-space search active
Cost-aware agentic planning making $c(a)$ a first-class term in the agent objective changes behaviour qualitatively, not just quantitatively active, from the verification-agent work
Binary/quantized accelerator co-design quantization schemes chosen jointly with the datapath beat schemes chosen for the model alone active, tied to ADRIN work
Verification-informed architecture designs that are cheap to verify are a distinguishable class, and the property is predictable from the RTL early

Tech Stack

Silicon & Verification

Machine Learning & AI

Languages, Frameworks & Cloud

Certifications

Certifications and coursework

Credential Issuer Why it's here
Hardware Security University of Maryland trojans, side channels, and PUFs — the attacks that live below the software threat model
VLSI Design L&T EduTech industry-framed digital design flow, end to end
Algorithms Specialization Stanford University the four-course track — divide & conquer, graphs, greedy/DP, NP-completeness
Agentic AI Foundations Associate 1Z0-1157-26 Oracle agent reasoning patterns, tool orchestration, MCP
AI Agents Course Hugging Face building and evaluating agents against real benchmarks
SQL AI Developer Associate DP-800 Microsoft vector search, embeddings and RAG pushed down into the database engine
Certified Data Engineer – Associate DEA-C01 Amazon Web Services pipelines, storage and orchestration — the layer the models actually run on
GitHub Foundations GH-900 GitHub the collaboration plumbing this whole profile is built on
Google Data Analytics Professional Google the eight-course analytics track

Hardware Security and the Algorithms specialization are the two that show up most in my actual work —
one because attacks find the layer you forgot to model, the other because complexity analysis is
the only honest way to argue about a design before you've built it.

Background

BS in Data Science — Indian Institute of Technology Madras
B.Tech in Electrical & Electronics Engineering — Amrita Vishwa Vidyapeetham

Two degrees, deliberately. One taught me to reason about data and uncertainty;
the other taught me what actually happens when electrons have to carry the answer.

GitHub Stats

3D contribution graph

Interface

Everything above this line is hand-authored SVG. No template, no generator, no screenshot.
The page is the portfolio piece — so here is the design system it runs on.

 🎨  The design system, and how the page is built

Colour tokens — one ramp from substrate to signal, with a violet axis reserved for anything quantum or probabilistic. substrate #030811 → die #071226 → well #0C1D3E → trace #1B3566, accented by signal #38BDF8 and signal-alt #22D3EE, text in ink #C8E0F8 and ink-dim #7FA6D4, probabilistic domains in quantum #A78BFA / #C084FC.

Type scale — three families, three jobs. Segoe UI for headings, JetBrains Mono for anything machine-adjacent, a math serif for equations. Section headers at 30px / 700 / +10 letter-spacing; annotations at 9–11px mono with lowered opacity, so they read as marginalia rather than content.

Motion, deliberately restrained — every animation is a 3–9 second loop with no easing spikes, so nothing competes for attention. Travelling pulses along circuit rails, a slow scan across banners, coverage bars that fill once and freeze, a Bloch vector that precesses. Motion signals aliveness, not urgency.

Layout — a fixed 1400px design width, 16–18px radii, 60px gutters, and a 10%-opacity dot grid on every panel so the page reads as one continuous surface rather than a stack of unrelated images.

Accessibility — every graphic carries alt text, nothing critical is colour-only, body contrast clears 7:1, and the page degrades to readable structured markdown if images fail.

Built from source, not by hand — every banner, divider and diagram is generated by parameterised Python in tools/, so the geometry is computed rather than eyeballed. Equations are typeset to SVG with MathJax and served light/dark through <picture>, because GitHub mangles double-dollar math inside collapsible blocks — it strips backslashes and eats underscore pairs as italics. The TeX source lives in tools/equations.json and in every image's alt. GitHub Actions regenerate the contribution graph nightly, and the dream with it.

Ask Me Anything

Got a question about any of this? World models, RTL, verification, the edge-AI work, or how this page is built — ask in the open.

Ask a question   Open threads

Answered in the open, so the next person with the same question can read it too.
A live, interactive Q&A widget is coming to the portfolio site.

Connect

Open to conversations about
world models and model-based RL  ·  AI accelerator architecture  ·  verification methodology
graph learning for EDA  ·  quantized and binary inference  ·  anything at the hardware/ML seam

GitHub LinkedIn Portfolio Email

Building efficient, intelligent systems where hardware, software, and learning meet.

Pinned Loading

  1. ai-chip-design-platform ai-chip-design-platform Public

    Multi-agent RISC-V verification and test-generation framework for AI-assisted RTL, ISS, compliance, coverage, and debug workflows.

    Python 12 4

  2. adaptive-llm-quantization adaptive-llm-quantization Public

    Agentic framework for hardware-aware, layer-wise LLM quantization using sequential optimization and on-device benchmarking.

    Python

  3. GRIT GRIT Public

    An AI-powered learning platform to help students prepare and practice for competitive exams like GATE, UPSC, and CAT.

    Python

  4. solar-digital-twin solar-digital-twin Public

    Cloud based solar digital twin simulation

    Python 4 1