Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions .github/ISSUE_TEMPLATE/bug_report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
---
name: Bug report
about: A documented flag, fix, or setup step doesn't actually work as described
title: ""
labels: bug
---

**What's documented, and where** (file + section/link):

**What you did**:

**What you expected** (per the docs):

**What actually happened**:

**Environment** — OS, `llama.cpp` build/commit, Open WebUI version, GPU: see [docs/compatibility.md](../../docs/compatibility.md) for what's already tested. If yours differs, say so explicitly:

**How you confirmed it's a real issue** (a command you ran, a log line, not just "seems off") — see [CONTRIBUTING.md](../../CONTRIBUTING.md) for why this matters:
5 changes: 5 additions & 0 deletions .github/ISSUE_TEMPLATE/config.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
blank_issues_enabled: true
contact_links:
- name: Security vulnerability
url: https://github.com/onkarbadve/localai/security/policy
about: Please report security issues privately per SECURITY.md, not as a public issue.
12 changes: 12 additions & 0 deletions .github/ISSUE_TEMPLATE/documentation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
---
name: Documentation improvement
about: A broken link, unclear explanation, or formatting issue in the docs
title: ""
labels: documentation
---

**File(s) affected**:

**What's wrong or unclear**:

**Suggested fix** (if you have one):
14 changes: 14 additions & 0 deletions .github/ISSUE_TEMPLATE/feature_request.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
name: Feature request
about: Suggest something for this setup — please read the scope note below first
title: ""
labels: enhancement
---

Per [CONTRIBUTING.md](../../CONTRIBUTING.md#whats-out-of-scope), this repo documents one person's actual hardware and choices, not a general-purpose product — feature requests for the setup itself are generally out of scope. This template is for the rare case that's still worth raising: e.g. a documented rough edge that has a concrete, verifiable fix, or a compatibility gap you've hit yourself.

**What you're running into**:

**What you'd want instead**:

**Have you verified this against your own setup** (not just a hypothetical)?
40 changes: 40 additions & 0 deletions .github/workflows/docs-lint.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
name: Docs Lint

# Documentation-only CI: this repo has no build or test suite (it's
# infrastructure/config, not a library — see AGENTS.md), so this workflow
# only validates markdown formatting and links. No build steps, no tests.
on:
push:
branches: [main]
paths: ["**/*.md"]
pull_request:
paths: ["**/*.md"]
workflow_dispatch:

permissions:
contents: read

jobs:
markdownlint:
name: Markdown formatting
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: DavidAnson/markdownlint-cli2-action@v19
with:
globs: |
**/*.md
!llama.cpp/**

link-check:
name: Markdown links
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: lycheeverse/lychee-action@v2
with:
args: >-
--config lychee.toml
--no-progress
"**/*.md"
fail: true
15 changes: 15 additions & 0 deletions .markdownlint.jsonc
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
{
// This repo writes prose as one paragraph per line (no hard-wrapping) and
// uses a couple of intentional structural choices (centered README header,
// compact tables) that trip the stricter default rules below without being
// actual problems. Disabled here rather than reformatting working content.
"default": true,
"MD013": false, // line-length — prose isn't hard-wrapped in this repo
"MD033": false, // no-inline-html — README's centered <div> header
"MD041": false, // first-line-heading — README's <div> precedes the H1
"MD022": false, // blanks-around-headings — pre-existing SETUP.md style
"MD032": false, // blanks-around-lists — pre-existing SETUP.md style
"MD024": false, // no-duplicate-heading — every docs/ file repeats "Related documents"
"MD060": false, // table-column-style — not yet applied repo-wide
"MD018": false // no-missing-space-atx — linkedin-post.md's real hashtags, not headings
}
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,14 @@

Version history of this repository and the setup it documents. This is a summary view — [JOURNAL.md](JOURNAL.md) is the detailed, dated source of truth; [adr/](adr/) explains the reasoning behind the decisions that stuck. Entries are grouped into rough milestones, not formal semver releases (there's no package here to version) — the numbers exist to give a sense of order and progress.

## v1.5 — Repository polish and CI (2026-07-31)

- README: tightened "Features" into "Key Takeaways" and "Repository Structure" into "Repository at a Glance" (added a quick-facts table) for first-screen scannability.
- Added `docs/compatibility.md` (tested OS/tool versions) and `docs/github-setup.md` (manual GitHub configuration checklist).
- Added a static PNG export of the architecture diagram (`docs/images/architecture.png`) alongside the existing Mermaid source.
- Added `.github/workflows/docs-lint.yml` — markdown formatting and link checks, no build/test steps — plus minimal issue templates (`.github/ISSUE_TEMPLATE/`).
- Terminology and cross-link consistency pass across every markdown file; no functional or technical content changed.

## v1.4 — Documentation overhaul (2026-07-30 onward)

- Restructured `README.md` around overview / features / architecture / quick start / repo structure / screenshots / benchmarks / documentation / roadmap.
Expand Down
2 changes: 2 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,8 @@ This repository is primarily a personal engineering notebook — a running recor

Read [AGENTS.md](AGENTS.md) before proposing documentation changes — it covers this repo's specific conventions (comment style in scripts, when to update `JOURNAL.md` vs. `docs/`, why content gets moved rather than deleted).

Maintainer-only GitHub configuration (repo description, topics, release strategy) is checklisted separately in [docs/github-setup.md](docs/github-setup.md) — it doesn't affect contributions, just repo settings.

## Code of conduct

Be respectful and constructive. This is a small, personal project maintained by one person in their spare time — response times may be slow.
25 changes: 25 additions & 0 deletions JOURNAL.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,31 @@ Older entries below predate this template and stay in their original free-form n

---

## 2026-07-31 — Final repository polish: discoverability, consistency, and CI

**Goal**: a documentation-only polish pass — improve first-time-visitor experience, terminology consistency, and navigation across the already-complete doc set, without adding new documentation or rewriting working content.

**Changes**:
- README: retitled "Features" → "Key Takeaways" (same evidence-backed bullets, cross-linked to the docs that back each claim) and "Repository Structure" → "Repository at a Glance" (added a compact quick-facts table above the existing directory tree).
- Fixed two Iris Xe wording inconsistencies (`adr/0004-vulkan-backend.md`, `docs/hardware.md`) to match the "Intel Iris Xe iGPU" phrasing used everywhere else.
- Exported the `docs/architecture.md` Mermaid diagram to a static PNG (`docs/images/architecture.png`, via `mmdc`) and referenced both versions in that doc, for viewers without Mermaid rendering.
- Added `docs/github-setup.md` — a reusable checklist for the GitHub-UI-only configuration (description, topics, social preview, homepage, release strategy, branch protection, discussions) that doesn't live in files.
- Added `docs/compatibility.md` — a tested-versions matrix for Fedora/Windows/Podman/Open WebUI/`llama.cpp`/Vulkan/Intel graphics drivers, with `TODO`s left wherever a version genuinely isn't recorded anywhere else in the repo, per the no-fabrication rule in `AGENTS.md`.
- Added `.github/workflows/docs-lint.yml` (markdownlint-cli2 + lychee link check, docs-only, no build/test steps), `.markdownlint.jsonc`, and `lychee.toml`.
- Added minimal issue templates (`.github/ISSUE_TEMPLATE/`: bug report, documentation, feature request + config.yml), the feature-request one explicitly pointing at `CONTRIBUTING.md`'s out-of-scope note given this repo documents one person's setup, not a product.
- Fixed small pre-existing markdown lint findings surfaced while configuring the lint job: a missing blank line around a fenced code block (`docs/troubleshooting.md`), and three bare URLs/email/unit-name false-positives wrapped correctly (`SECURITY.md`, `SETUP.md`, `linkedin-post.md`).
- Cross-linked the two new docs into README's Documentation table, `docs/hardware.md`, `docs/troubleshooting.md`, and `CONTRIBUTING.md`'s "Related documents"/workflow sections.

**Results**: `npx markdownlint-cli2 "**/*.md"` and a custom internal-link checker both run clean (the only "broken" links are the two pre-existing, expected ones into the gitignored `llama.cpp/` checkout — see `docs/hardware.md#why-models-and-llamacpp-arent-in-git`). `codespell` found nothing beyond one confirmed false positive (`</nothink>`, a literal template tag). README grew from 167 to 179 lines — within the "don't significantly increase length" constraint for this task.

**Problems**: none blocking. The task brief referenced README sections ("Key Takeaways", "Repository at a Glance") that didn't exist under those exact names yet — treated as a rename/tighten of the closest existing sections ("Features", "Repository Structure") rather than new, separately-maintained content, to avoid duplication.

**Lessons**: this repo's existing documentation was already unusually consistent (single terminology convention, dense cross-linking, no fabricated numbers) — most of the "audit" phases of this task confirmed a clean bill of health rather than finding real problems, which is itself worth recording so a future pass doesn't re-litigate the same ground.

**Next steps**: none opened by this pass. `docs/compatibility.md`'s `TODO` rows (Windows edition/build, Podman version, `llama.cpp`/Vulkan/Mesa versions on Fedora) are real gaps — worth filling in next time any of those get touched for an unrelated reason, not urgent enough to chase down standalone.

---

## 2026-07-30 — Documented the Podman-over-Docker rationale (retroactively, no re-litigation)

User asked why Podman was chosen for Open WebUI/Open Terminal over Docker; no prior journal entry recorded an explicit decision, so it was never actually debated in-session - Podman was simply the starting choice both container scripts were built on (`start-open-webui.sh`, `start-open-terminal.sh`, see SETUP.md). Capturing the reasoning now so it isn't lost:
Expand Down
22 changes: 17 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,14 +39,14 @@ This repo is both a working setup and its own documentation: every bug, workarou
- Enterprise AI infrastructure or multi-tenant deployment
- Distributed/clustered serving across multiple machines

## Features
## Key Takeaways

- **Zero cloud dependency** — every model runs locally on integrated graphics, nothing leaves the machine.
- **Vulkan iGPU offload** — full-layer GPU offload on an Intel Iris Xe iGPU, no discrete card required.
- **Vulkan iGPU offload** — full-layer GPU offload on an Intel Iris Xe iGPU, no discrete card required, sustaining ~9–9.8 tok/s (full numbers in [`docs/benchmarks.md`](docs/benchmarks.md)).
- **Dual-model chat** — a production agent-mode model and a separate blunt/uncensored model, both registered in Open WebUI, run mutually exclusively without reconfiguration.
- **Agent tooling** — real tool-calling (Builtin Tools) and a shell/file Integration (Open Terminal), verified end-to-end through Open WebUI, not just curl.
- **Agent tooling** — real tool-calling (Builtin Tools) and a shell/file Integration (Open Terminal), verified end-to-end through Open WebUI in 9/9 real tests, not just curl.
- **Remote access** — reachable from a phone over a Tailscale VPN overlay, no port-forwarding or public exposure.
- **Documented crash hardening** — local patches and mitigations for a real iGPU fence-timeout bug, turning silent crashes into clean, recoverable errors.
- **Documented crash hardening** — local patches and mitigations for a real iGPU fence-timeout bug, turning silent crashes into clean, recoverable errors (see [`docs/troubleshooting.md`](docs/troubleshooting.md)).

## Architecture

Expand Down Expand Up @@ -81,7 +81,17 @@ This is the simplified shape of it. For the full diagram — both `llama-server`

Models aren't checked into this repo (multi-GB GGUF files). Building `llama.cpp` from source, placing model weights, Tailscale setup, and every flag's reasoning are in [`SETUP.md`](SETUP.md) — start there for anything beyond running an already-built setup.

## Repository Structure
## Repository at a Glance

| | |
|---|---|
| **Hardware** | Intel i5-12500H · Intel Iris Xe iGPU · 16GB RAM — no discrete GPU |
| **OS** | Fedora 44 (daily driver) · Windows (origin, preserved) |
| **Inference** | `llama.cpp`, Vulkan backend, full-layer GPU offload |
| **Chat frontend** | Open WebUI + Open Terminal, rootless Podman |
| **Remote access** | Tailscale VPN overlay |
| **License** | MIT |
| **Status** | Active daily driver ([`JOURNAL.md`](JOURNAL.md)) |

```text
LocalAI/
Expand Down Expand Up @@ -137,6 +147,8 @@ Full numbers, multi-turn cache-reuse data, and agent-mode round-trip timings: [`
| [`docs/troubleshooting.md`](docs/troubleshooting.md) | Structured Problem/Cause/Solution/Verification writeups |
| [`docs/lessons-learned.md`](docs/lessons-learned.md) | Practical conclusions from real experimentation, by topic |
| [`docs/roadmap.md`](docs/roadmap.md) | Completed / upcoming / future-idea work, in more detail than below |
| [`docs/compatibility.md`](docs/compatibility.md) | Tested versions of every OS/tool in the stack |
| [`docs/github-setup.md`](docs/github-setup.md) | Manual GitHub configuration checklist (topics, social preview, releases) |
| [`adr/`](adr/) | Architecture Decision Records — why the stable, load-bearing choices were made |
| [`JOURNAL.md`](JOURNAL.md) | Dated running log — newest entries first, the historical source of truth |
| [`CHANGELOG.md`](CHANGELOG.md) | Version history of this repository itself |
Expand Down
2 changes: 1 addition & 1 deletion SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ Relevant components, for context on what a report might touch:

If you find a real security issue in **this repository's own scripts or documented configuration** (for example: a script that would expose a service more broadly than documented, an API key handled unsafely, or a documented mitigation that's actually ineffective), please report it privately rather than opening a public issue:

**onkarbadve@gmail.com**
**<onkarbadve@gmail.com>**

Include:
- What you found and why it's a security issue (not just a bug).
Expand Down
2 changes: 1 addition & 1 deletion SETUP.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,7 +123,7 @@ Rootless Podman container (`ghcr.io/open-webui/open-webui:main`), point at which
- Data persisted in a named Podman volume (`open-webui-data`), not bind-mounted.
- Script is idempotent: first run creates the container, later runs just `podman start` the existing one.

UI: http://localhost:3000 — first visit creates the local admin account. Switching which llama.cpp model is active just means restarting the underlying `start-*.sh` script for the port you want live; Open WebUI already has both 8080 and 8081 registered as connections (see above), so it reconnects automatically rather than needing reconfiguration.
UI: <http://localhost:3000> — first visit creates the local admin account. Switching which llama.cpp model is active just means restarting the underlying `start-*.sh` script for the port you want live; Open WebUI already has both 8080 and 8081 registered as connections (see above), so it reconnects automatically rather than needing reconfiguration.

Not yet ported to Windows (no Podman/Docker there currently).

Expand Down
2 changes: 1 addition & 1 deletion adr/0004-vulkan-backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@

## Context

There is no discrete GPU anywhere in this setup — the only GPU-class compute device available is the Intel Iris Xe integrated GPU (see [docs/hardware.md](../docs/hardware.md)). Running models at usable speed on a 16GB machine requires real GPU offload rather than falling back to CPU-only inference.
There is no discrete GPU anywhere in this setup — the only GPU-class compute device available is the Intel Iris Xe iGPU (see [docs/hardware.md](../docs/hardware.md)). Running models at usable speed on a 16GB machine requires real GPU offload rather than falling back to CPU-only inference.

## Decision

Expand Down
4 changes: 4 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,10 @@ flowchart LR
QwenU -->|Vulkan| GPU
```

Static export of the diagram above, for viewers without Mermaid rendering: [`images/architecture.png`](images/architecture.png).

![Architecture diagram](images/architecture.png)

## Components

- **`llama-server` (×2 registered, 1 running at a time)** — the inference engine, built from `llama.cpp` source with the Vulkan backend. Production (`start-qwen3.sh`, port `8080`) and an uncensored/blunt-mode variant (`start-qwen3-uncensored.sh`, port `8081`) both register as connections in Open WebUI, but only one process runs at once — combined Vulkan memory doesn't fit both simultaneously at their current context sizes (see [SETUP.md](../SETUP.md#qwen3-4b-instruct-2507-heretic-av2--uncensored-bluntdirect-assistant-start-qwen3-uncensoredsh-fedora-only)). Whichever port is live just shows up in Open WebUI's model picker.
Expand Down
27 changes: 27 additions & 0 deletions docs/compatibility.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Compatibility Matrix

What this setup has actually been run against, not a general support statement. Sourced from [SETUP.md](../SETUP.md), [JOURNAL.md](../JOURNAL.md), and [hardware.md](hardware.md). Per [AGENTS.md](../AGENTS.md#documentation-standards), unmeasured values are left as `TODO` rather than guessed — if you run this setup on a version not listed here and confirm it works (or doesn't), see [CONTRIBUTING.md](../CONTRIBUTING.md) for how to report it.

| Component | Tested version | Notes |
|---|---|---|
| Fedora | 44 | KDE Plasma spin, current daily driver. See [hardware.md](hardware.md). |
| Windows | TODO — edition/build not recorded | Origin OS, preserved on a separate NTFS partition. See [hardware.md](hardware.md#windows-vs-fedora). |
| Podman | TODO — Fedora 44's default repo version, exact version not recorded | Rootless, no daemon. See [architecture.md](architecture.md#design-decisions-worth-calling-out). |
| Open WebUI | `ghcr.io/open-webui/open-webui:main` (floating tag, not pinned) | Deployed via [`start-open-webui.sh`](../start-open-webui.sh). A `:main` tag means "whatever's current" — pin to a release tag if reproducibility matters more than staying current. |
| Open Terminal | `ghcr.io/open-webui/open-terminal:slim` (floating tag, not pinned) | Deployed via [`start-open-terminal.sh`](../start-open-terminal.sh). Same floating-tag caveat as above. |
| `llama.cpp` (Fedora) | Built from source, Vulkan backend (`GGML_VULKAN=ON`) | Commit/tag in use not currently recorded — see [hardware.md](hardware.md#why-models-and-llamacpp-arent-in-git). |
| `llama.cpp` (Windows) | Precompiled Vulkan release, upgraded b9305 → b10107 | See [benchmarks.md](benchmarks.md#build-upgrade-impact-windows-b9305--b10107) for the re-verification after upgrade. |
| Vulkan | TODO — SDK/driver version not recorded | Backend confirmed working via GPU-resident inference (low RSS) on both OSes; see [adr/0004-vulkan-backend.md](../adr/0004-vulkan-backend.md). |
| Intel graphics driver | TODO — Mesa/`i915` version not recorded | Relevant to the fence-timeout bug in [troubleshooting.md](troubleshooting.md#i915-igpu-fence-timeout--gpu-hangs); worth recording if that bug is ever revisited upstream. |
| Tailscale | 1.98.8 (Fedora, installed from Fedora's own repos) | See [SETUP.md](../SETUP.md) for the firewalld zone configuration needed alongside it. |

## Why this exists

Every fix and workaround in [troubleshooting.md](troubleshooting.md) is tied to a specific version combination, even where that version isn't pinned down precisely yet — a driver update, a Vulkan SDK bump, or an Open WebUI release could change or resolve any of them. This table exists so a future reader (including future-self) can tell whether a fix is still expected to apply, and so gaps (the `TODO`s above) are visible instead of silently assumed.

## Related documents

- [hardware.md](hardware.md) — the machine these versions run on
- [troubleshooting.md](troubleshooting.md) — issues tied to specific versions above
- [SETUP.md](../SETUP.md) — full flag-by-flag configuration
- [AGENTS.md](../AGENTS.md#documentation-standards) — why unmeasured values stay `TODO` instead of guessed
Loading
Loading