Skip to content

Repository files navigation

dl-neural-networks

A dense network on MNIST and Fashion-MNIST: the same architecture on two datasets, and what the accuracy gap between them says about the limits of fully-connected layers on images.

The gap is usually explained as "flattening loses spatial information", which sounds like a matter of degree. It is exact and total, and the project demonstrates that by construction: for any dense network and any fixed permutation of the pixels, there is a dense network producing bit-identical outputs on the shuffled images. That is the concrete reason convolution exists.

Standard library at runtime. 76 tests.

Skills demonstrated

Deep learning theory — a constructive proof of permutation equivalence for fully-connected layers, the architectural reason convolution is not interchangeable with it, weight-sharing over a neighbourhood as the property that breaks under permutation

ML engineering — controlled dataset pair holding resolution, class count and split sizes fixed so the accuracy gap is attributable to the images; error-rate ratio reported beside the raw gap; numerically stable softmax

Systems / data — IDX binary format parsed from the specification with struct and gzip, magic-number type checking that rejects a float file rather than misreading it as uint8

Software engineering — a forward-pass network in the standard library so the central claim is a unit test rather than a training run; frozen dataclasses with validation in __post_init__; 76 tests covering ragged weights, mismatched layers, truncated files and permutation round-trips

Tooling — ruff, pre-commit, CI matrix on 3.10/3.11/3.12 with a lint job

What the run looks like

$ python examples/permutation_blindness.py
==========================================================================
1. The same architecture, two datasets

mnist     0.9800
fashion   0.8800
gap +0.1000   error rate 6.0x higher

==========================================================================
2. Why flattening costs nothing on one and a great deal on the other

dense network: 101,770 parameters, 784 inputs
128 inputs, 784 pixels shuffled by a fixed permutation
  outputs identical to the last bit
  predictions changed: 0

==========================================================================
3. The same claim on one concrete input

  original network on the image       ['0.094299', '0.031709', '0.017338', '0.118151'] ...
  permuted network on shuffled pixels ['0.094299', '0.031709', '0.017338', '0.118151'] ...
  identical: True

The six decisions worth discussing

1. The permutation result is constructive, not statistical. Fix any permutation P of the 784 inputs. Reordering the columns of the first weight matrix by P gives a network N' with N'(Px) == N(x) for every x. No arithmetic is redone — the same products are summed, in a different order — so the equality is exact rather than approximate. equivalence_error reports the largest observed difference, and it is 0.0.

2. The demonstration uses untrained weights, deliberately. The property is architectural: it holds for every possible solution, not for the one gradient descent happened to find. Showing it with random weights makes that explicit, and removes any suspicion that it is an artefact of a particular training run.

3. The dataset pair is the control. MNIST and Fashion-MNIST have identical resolution, identical class counts and identical split sizes. Only the images differ, so the accuracy gap is a property of the images — which is what licenses reading it as "how much this dataset depends on spatial structure". A comparison against CIFAR-10 would confound resolution, colour and class count all at once.

4. The error ratio is reported next to the gap. 0.98 to 0.88 is ten points and six times the error rate. The second number is the one that matters once an error budget exists, and the first is the one that gets quoted.

5. The IDX reader checks the magic number's type code. An IDX file of 32-bit floats read as uint8 yields an image that looks like noise and raises nothing at all. Fifteen lines of struct replace a dependency and make the failure loud.

6. Permutation distinguishes gather from scatter. order[i] is the source position landing at i. The inverse is a different array, and confusing the two produces a permutation that looks correct and is the wrong way round — so inverse exists and a test composes them back to the identity.

Why the network is written out

mlp.py is a forward pass in about a hundred lines of standard library. That makes the central claim testable in milliseconds with no TensorFlow, no GPU and no download — test_permute.py checks the equivalence on hand-built two-pixel examples, on random networks, and on a three-layer stack.

python -m denselimits --train runs the real comparison: dense and convolutional networks, on both datasets, with and without the pixel permutation. The dense accuracies are unchanged by the shuffle and the convolutional ones collapse.

Design

src/denselimits/
  mlp.py         Dense/MLP forward pass, stable softmax
  permute.py     Permutation, permute_network, the equivalence check
  idx.py         the IDX binary format, via struct and gzip
  datasets.py    the controlled MNIST / Fashion-MNIST pair
  metrics.py     accuracy, per-class scores, top confusions
  experiment.py  the dataset gap and the permutation result
  models.py      the only module that imports TensorFlow
  cli.py         argument parsing

The boundary that matters: nothing outside models.py imports TensorFlow, and nothing outside idx.py knows how the files are encoded.

Usage

python3.12 -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
pytest -q
python examples/permutation_blindness.py
python -m denselimits
python -m denselimits --pixels 256 --hidden 64 --samples 32

Training for real:

pip install -r requirements.txt
python -m denselimits --train --epochs 10

As a library:

from denselimits import Permutation, permute_network, random_network

network = random_network(784, 128, 10, seed=0)
permutation = Permutation.random(784, seed=0)
permuted = permute_network(network, permutation)

image = [0.0] * 784
assert network.forward(image) == permuted.forward(permutation.apply(image))

Dataset

MNIST — 70,000 28×28 grayscale handwritten digits, 10 classes. Fashion-MNIST — 70,000 28×28 grayscale garment images, 10 classes, a drop-in replacement with identical shape.

Neither is included; about 30 MB each. Keras downloads and caches both:

from tensorflow import keras

(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_data()
(x_train, y_train), (x_test, y_test) = keras.datasets.fashion_mnist.load_data()

Or download the IDX files directly and read them with no TensorFlow at all:

curl -O https://storage.googleapis.com/cvdf-datasets/mnist/train-images-idx3-ubyte.gz
curl -O https://storage.googleapis.com/cvdf-datasets/mnist/train-labels-idx1-ubyte.gz
from denselimits import read_images, read_labels

images = read_images("train-images-idx3-ubyte.gz", limit=1000)
labels = read_labels("train-labels-idx1-ubyte.gz", limit=1000)
print(images.render(0))

Fashion-MNIST ships the same four files in the same format from https://github.com/zalandoresearch/fashion-mnist.

Nothing needs downloading to reproduce the permutation result.

Scope

  • The forward pass is implemented; training is not. The equivalence is a property of the architecture, so it needs no optimiser. Training the real comparison is delegated to Keras behind --train.
  • Accuracies in the dataset gap are the documented typical values for a 784-128-10 network, used for orientation. gap(accuracies=...) takes measured numbers, and --train produces them.
  • One permutation per run. The result holds for every permutation by construction, so sampling more of them would add runtime and no information.
  • No convolutional implementation in the standard library. The control for the permutation test is Keras-side. Writing a conv forward pass would not strengthen the argument, which is about what dense layers cannot do.
  • The gap is attributed to spatial structure, not decomposed. Separating texture from silhouette from intra-class variance is a different experiment needing per-class ablations on the images themselves.

License

MIT — see LICENSE.

About

A dense network on MNIST and Fashion-MNIST, with a constructive proof that fully-connected layers are exactly invariant to pixel permutation — the concrete reason convolution exists. Forward pass in the standard library, 76 tests.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages