Skip to content

Latest commit

Β 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Formalizing Latent Thoughts:
Four Axioms of Thought Representation in LLMs

arXiv Hugging Face Paper Project Website GitHub Stars X Coverage LinkedIn Coverage WisPaper Coverage Daily AI Wire Coverage MIT License


Four axioms of thought representation in LLMs


News

[2026.06.30] Check out the WisPaper blog and Daily AI Wire article covering our research!

[2026.06.29] Our paper is featured as πŸ€— HuggingFace #1 Paper of the Day!

[2026.06.29] Covered by AK on X / HuggingFace Daily Papers.

[2026.06.29] Paper launch: X announcement and LinkedIn post.

[2026.05.07] We have released our paper on arXiv!


Abstract

We introduce an axiomatic evaluation framework for latent thought representations in LLMs, comprising metrics that are independent of downstream benchmark scores and reveal representational failures that benchmark accuracy masks. Existing evaluations conflate representation quality with model capacity. Therefore, failures cannot be attributed to the representation rather than to the model that processes it. We formalize four functional axioms β€” Causality, Minimality, Separability, and Stability β€” and define a quantitative measure for each, computed directly on the representation independently of downstream accuracy. We audit open-weight LLMs across 23 reasoning tasks (e.g., Spatial Reasoning, Factual QA). We find that no candidate satisfies all four axioms simultaneously, that the representations distinguish task type reliably but cannot distinguish between two questions within the same task, and that the representations encode little information beyond what is already present in the input embedding. The failure is consistent across dense, reasoning-distilled, and RL-trained model families, indicating that the gap is structural rather than a property of model size or training procedure.


Demo

Screencast.from.2026-06-28.18-52-08.mp4


Requirements

Python 3.12, CUDA 12.6, GCC 12.3.1.

Install uv, then:

uv pip install torch
uv sync

Optional flash-attention support (set CUDA_HOME and add to PATH first):

uv pip install psutil ninja packaging einops setuptools wheel
uv pip install flash-attn --no-build-isolation

Setup

Copy .env.example to .env and fill in your credentials:

cp .env.example .env

Download benchmarks before running the pipeline:

uv run python -m scripts.download.download_bbeh

Pipeline

All scripts use Hydra and must be run from the project root with uv run python -m. Phases must run in order: 1 β†’ 2 β†’ 3 β†’ 4, then minimality, causality, and DCS can run in any order after Phase 4.

Phase 1 β€” LLM Data Generation

Generates LLM responses and first-token prefill hidden states for all 23 BBEH tasks.

# Llama-3.1-8B (default)
uv run python -m scripts.llm_data

# Other models β€” one config per model
uv run python -m scripts.llm_data --config-name=llm_data_70b
uv run python -m scripts.llm_data --config-name=llm_data_deepseek_r1_32b
uv run python -m scripts.llm_data --config-name=llm_data_skywork_or1_32b
uv run python -m scripts.llm_data --config-name=llm_data_gpt_oss_20b

# Quick smoke-test (10 examples, no wandb)
uv run python -m scripts.llm_data loader=dev wandb.use_wandb=false

Outputs land in outputs/llm_data_<model>/.

Phase 2 β€” Discriminator Index

Builds positive/negative pair indices from Phase 1 outputs. Point llm_data_output_dir at the Phase 1 output.

uv run python -m scripts.disc_index \
    llm_data_output_dir=outputs/llm_data_8B

Outputs land in outputs/disc_index_output/.

Phase 3 β€” Discriminator Training Data

Pre-generates cached TR vectors for all thought representation types. Must match the source model used in Phase 1.

# Llama-8B (hidden size 4096)
uv run python -m scripts.disc_data \
    llm_data_output_dir=outputs/llm_data_8B \
    disc_index_output_dir=outputs/disc_index_output

# Llama-70B (hidden size 8192)
uv run python -m scripts.disc_data \
    --config-name=disc_data \
    base_llm=llama_70b \
    llm_data_output_dir=outputs/llm_data_70b \
    discriminator.other_vector_dim=8192

Outputs land in outputs/disc_data_<model>/.

Phase 4 β€” Discriminator Training (Separability)

Trains one discriminator per TR type. Key overrides: tr_type, think_steps (for soft/latent thinking), source_hidden_size (must match source model).

# Example: last_input_token on Llama-8B
uv run python -m scripts.disc_trainer \
    tr_type=last_input_token \
    disc_data_output_dir=outputs/disc_data_8B

# Example: soft_thinking at 128 steps
uv run python -m scripts.disc_trainer \
    tr_type=soft_thinking \
    think_steps=128 \
    disc_data_output_dir=outputs/disc_data_8B

# Example: Llama-70B (override hidden size and layer count)
uv run python -m scripts.disc_trainer \
    tr_type=last_input_token \
    source_hidden_size=8192 \
    source_num_layers=81 \
    disc_data_output_dir=outputs/disc_data_70b

Outputs land in outputs/discriminator/.

Minimality Probe

Trains two probe families per TR type (both required for the IB-residual gap Delta_IB):

# Probe 1: predict Y from T  (ygt)
uv run python -m scripts.minimality_trainer \
    --config-name minimality_train_output \
    tr_type=last_input_token \
    llm_data_output_dir=outputs/llm_data_8B \
    tr_data_output_dir=outputs/disc_data_8B

# Probe 2: predict X from (Y, T)  (xgyt)
uv run python -m scripts.minimality_trainer \
    --config-name minimality_train_xgyt \
    tr_type=last_input_token \
    llm_data_output_dir=outputs/llm_data_8B \
    tr_data_output_dir=outputs/disc_data_8B

For soft/latent thinking, also pass think_steps=128. For non-8B models, also pass probe.vector_dim=<hidden_size> (70B β†’ 8192, 32B β†’ 5120).

Generate qualitative output samples:

uv run python -m scripts.minimality_samples

Causality Evaluation

KL-substitution evaluation. Requires a trained discriminator from Phase 4. Use tile_to_length=128 to match the paper's tiling protocol.

uv run python -m scripts.causality_eval \
    tr_type=last_input_token \
    disc_dir=outputs/discriminator/llama_8b/last_input_token \
    tr_data_dir=outputs/disc_data_8B \
    tile_to_length=128

To use the minimality projection (appendix ablation), add:

    proj_source=minimality_output \
    min_proj_run_label=<RUN_LABEL> \
    min_proj_root=outputs/min_prob_output

DCS β€” Distributional Consistency Score

Requires a trained discriminator from Phase 4.

uv run python -m scripts.dcs_eval \
    tr_type=last_input_token \
    disc_dir=outputs/discriminator/llama_8b/last_input_token \
    tr_data_dir=outputs/disc_data_8B

uv run python -m scripts.dcs_d_eval \
    tr_type=last_input_token \
    tr_data_dir=outputs/disc_data_8B

Thought Representation Types

Code Description
last_input_token Hidden states from last prefill token, all layers
last_input_hidden_state Single hidden state at last prefill token
soft_thinking Iterative soft-thinking representations
soft_thinking_noise Soft-thinking with noise injection
latent_thinking Latent iterative representations
embedding_no_pooling Per-beam text embeddings (no pooling)
embedding_pooling Pooled text embeddings
input_embedding Input-only text embeddings
random_vector Random baseline

Running Tests

uv run pytest tests/
uv run pytest tests/test_phase2.py::TestName

Citation

@misc{seddik2026formalizinglatentthoughtsaxioms,
  title         = {Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs},
  author        = {Fahd Seddik and Fatemeh Fard},
  year          = {2026},
  eprint        = {2606.27378},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2606.27378}
}

License

This project is released under the MIT License.

About

Official implementation of Formalizing Latent Thoughts: Four Axioms for Evaluating Thought Representations in LLMs

Topics

Resources

Stars

7 stars

Watchers

2 watching

Forks

Contributors

Languages