Skip to content
View AnnasMazhar's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report AnnasMazhar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AnnasMazhar/README.md

Annas Mazhar

Data engineer. I build data platforms on AWS and, outside work hours, small autonomous systems and research tooling — mostly Python, always with a bias for things that can be checked rather than believed.

The rule I build to: if a tool makes a claim, it should be able to prove it. Each repo below ships the tests, an evidence file, and the failure modes it does not hide.


Current work

factorproof — factor research that refuses to promote noise

Purged walk-forward cross-validation with embargo, IC decay curves, Benjamini-Hochberg FDR across the whole factor family, deflated Sharpe for the number of trials actually run — and a promotion gate a factor must satisfy before it is allowed into a portfolio.

$ factor-lab list
Name                   Category           Description
mom_20                 momentum           20-day log-return (short-term momentum)
rev_5                  reversal           Negated 5-day log-return (short-term reversal)
vol_20                 volatility         20-day realised volatility, annualised
rsi_14                 oscillator         14-day Relative Strength Index (Wilder smoothing)

107 tests · MIT

replayproof — offline regression testing for LLM agents

Record an agent run once, re-run it forever with zero API keys. Assert tool-call contracts, gate the token cost of a change, and diff behaviour when you swap the model underneath.

$ agenteval replay --run examples/recordings/sample_run.jsonl --mode strict
$ agenteval drift  --baseline base_result.json --current new_result.json

136 tests · MIT

fitsproof — local inference that proves it fits

A calibrated cost model and an enforced memory budget in front of a from-scratch NumPy inference engine (KV cache, quantisation). When the model does not fit, it says so loudly instead of OOM-ing silently at token 400.

$ fitsproof probe
Probing machine...
  bandwidth:  6.69 GB/s
  gemm:       314.31 GFLOPS
  RAM:        33.55 GB
  VRAM:       0.00 GB

148 tests · MIT

warehouse-mcp — least-privilege SQL for agents (not yet public)

An MCP server that gives an AI agent read access to a warehouse through a SQL-AST-enforced column policy, row filters, query budgets, redaction and an append-only audit trail. In progress.


How these are built

Each repo went through the same iteration protocol — research → implement → evaluate → adversarial review → mutation testing — with the evidence committed next to the code (EVIDENCE.md, docs/ADVERSARIAL_REVIEW.md, reports/).

The test counts above are not self-reported. They come from a fresh clone, a clean virtualenv, pip install -e '.[dev]', pytest and ruff — the same commands in each README — and run on CI for Python 3.11 and 3.12.

Where something is unfinished or wrong, the README says so. A repo that hides its gaps is a demo; one that states them is a piece of work.


Elsewhere

Pinned Loading

  1. PyPerfOptimizer PyPerfOptimizer Public

    Make Python faster. Automatically. Profile → Detect → Transform → Verify.

    Python

  2. pyspark_mcp pyspark_mcp Public

    SQL To Pyspark converter

    Python

  3. pytest-duckdb pytest-duckdb Public

    pytest plugin for SQL pipeline testing with DuckDB — load fixtures, run queries, snapshot results. Zero config.

    Python

  4. systemd-doctor systemd-doctor Public

    Fleet health CLI for systemd — traffic-light status, overdue timers, crash loops, security scoring. Zero dependencies.

    Python

  5. factorproof factorproof Public

    Evidence-gated factor research: purged walk-forward CV, IC decay, Benjamini-Hochberg FDR and deflated Sharpe — with a promotion gate that refuses to promote noise.

    Python

  6. fitsproof fitsproof Public

    Local LLM inference that proves it fits: calibrated cost model, enforced memory budget, loud degradation instead of silent OOM, and a from-scratch NumPy engine with KV cache and quantisation.

    Python