Skip to content

fix: save the config of diffusers models on export - #2119

Merged
changwangss merged 1 commit into
intel:mainfrom
Ar4ikov:fix-diffusers-config-export
Sep 16, 2026
Merged

changwangss merged 1 commit into
intel:mainfrom
Ar4ikov:fix-diffusers-config-export

Conversation

@Ar4ikov

@Ar4ikov Ar4ikov commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Description

quantize_and_save on a diffusers ModelMixin writes every safetensors shard and then dies before writing any config, leaving a checkpoint nothing can load.

AttributeError: 'FrozenDict' object has no attribute 'save_pretrained'
  File "auto_round/export/utils.py", line 48, in _save_model_configs
    model.config.save_pretrained(save_dir)

A diffusers model's .config is a diffusers.configuration_utils.FrozenDict, which has no save_pretrained; ModelMixin.save_config is the equivalent writer. The diffusion path never hits this because diffusion_load_model patches a save_pretrained onto the pipeline and component configs (auto_round/utils/model.py, around the config_save_pretrained partials). So the crash is reserved for callers who load a transformer themselves and pass it to AutoRound, and what they get is shards on disk with no config.json and no quantization_config.

One detail that made the obvious fix wrong: calling model.save_config(save_dir) alone writes a config.json without quantization_config, because the exporter sets that on the config object (export_to_autoround/export.py: model.config.quantization_config = quantization_config) and save_config only serializes the config's own dict. I verified that directly — save_config after setting the attribute produces a config.json whose keys do not include quantization_config. So the helper writes the config with diffusers' own writer and then merges quantization_config back in.

The same model.config.save_pretrained(...) call exists at two more sites, so all three now go through one helper: formats.py (fake format, meta-device branch) and export_to_llmcompressor/export.py. I only exercised the auto_round format path with a diffusers model; the other two are the same substitution, unverified against a diffusers model. export_to_mlx has the same call and is left alone — a diffusers transformer is not an MLX target.

Reproduction

Against main at 417d87d, with torch 2.11.0+cu128 / diffusers 0.40.0.dev0 / transformers 5.14.1:

import torch
from diffusers import SD3Transformer2DModel
from auto_round import AutoRound

SD3Transformer2DModel(
    sample_size=8, patch_size=2, in_channels=4, num_layers=1,
    attention_head_dim=32, num_attention_heads=2, joint_attention_dim=64,
    caption_projection_dim=64, pooled_projection_dim=64, out_channels=4,
).save_pretrained("/tmp/tiny_sd3")
model = SD3Transformer2DModel.from_pretrained("/tmp/tiny_sd3", torch_dtype=torch.bfloat16)

class _StubTokenizer:  # see note below
    pass

AutoRound(
    model=model, tokenizer=_StubTokenizer(), scheme="W4A16", group_size=32, sym=True,
    iters=0, disable_opt_rtn=True, to_quant_block_names="transformer_blocks", batch_size=1,
).quantize_and_save("/tmp/out", format="auto_round", inplace=True)

Before: AttributeError as above, and the only thing on disk is w4g32/model.safetensors.

After: no error, and

w4g32/config.json
w4g32/model.safetensors
w4g32/quantization_config.json

with config.json carrying _class_name: SD3Transformer2DModel, _diffusers_version, and a quantization_config (quant_method: auto-round, bits: 4, group_size: 32, packing_format: auto_round:auto_gptq).

Two honest notes about that snippet:

  • The _StubTokenizer is not part of this bug. On current main, ModelContext raises ValueError("A tokenizer must be set for non-str model input") for any model object, including data-free RTN, because BaseOrchestrator.__init__ passes need_calib=self.need_calib (still the class default True) before self.need_calib = self._check_need_calib() runs later in the same __init__. That looks like a separate ordering bug; happy to file it separately.
  • The shards are still named model-*.safetensors, while diffusers looks for diffusion_pytorch_model*. This PR does not touch naming — rename_weights_files already exists for that and is used by the diffusion path.

Test

test/test_cpu/export/test_export.py::test_save_model_writes_diffusers_config builds a tiny SD3 transformer in-process (no download) and calls save_model(..., immediate_saving=True), the path that crashed.

# with the fix
1 passed, 26 deselected in 0.43s

# with auto_round/ reverted, test kept
auto_round/export/utils.py:48: AttributeError
1 failed, 26 deselected in 0.33s

Where this came from

Quantizing MiniMax-H3's 33B video/audio DiT (MiniMaxH3Transformer3DModel) to W4A16. The transformer is loaded directly rather than through a pipeline, so the export died after writing all 14 shards; the config had to be reconstructed by hand from the tensors on disk. Result and write-up: https://huggingface.co/Ar4ikov/MiniMax-H3-transformer-W4A16-RTN

Type of Change

Bug fix

Related Issues

No issue filed — the failure is deterministic and the fix is small. Happy to open one if you prefer it tracked.

Checklist Before Submitting

  • My code has been tested locally. black / isort / ruff / codespell clean on the touched files. On the export subset (autogptq_format, autoround_format, awq_format, immediate_saving) the patched tree gives the same result as unpatched main in the same environment: 5 failed, 1 passed either way, plus the new test passing. All five failures are check_compressed_tensors_supported (ImportError: Please install compressed-tensors ... / SystemExit: -1) — a missing optional dependency here, which also means the llm_compressor export path is not covered by my local run.
  • Documentation has been updated as needed. No docs describe this path.
  • New or updated tests are included where applicable.
  • The CUDA CI has passed. I cannot trigger /azp run Unit-Test-CUDA-AutoRound on this repo; please kick it off.

@WeiweiZhang1, you last touched this config-saving code in #1810, and @mengniwang95 as the author of diffusion model saving (#1519), where the save_pretrained-on-the-config patch that hides this bug comes from.

@wenhuach21
wenhuach21 requested a review from changwangss August 4, 2026 09:32
Comment thread auto_round/export/export_to_llmcompressor/export.py
Comment thread auto_round/export/utils.py
@wenhuach21 wenhuach21 added this to the 0.16.0 milestone Sep 3, 2026
@wenhuach21

Copy link
Copy Markdown
Contributor

@changwangss could you help fix this issue if the author is not free?

The export path ended with `model.config.save_pretrained(save_dir)`. A diffusers
`ModelMixin` keeps its config in a `FrozenDict`, which has no such method, so
`quantize_and_save` on one wrote every shard and then died with an
AttributeError: no config.json, no quantization_config, nothing that can be
loaded back.

The diffusion path escapes this only because `diffusion_load_model` patches a
`save_pretrained` onto the pipeline config, so anyone who loads a transformer
themselves and hands it to AutoRound gets the half-written checkpoint.

Route all three config-saving sites through one helper that uses
`ModelMixin.save_config` for diffusers models and merges back the
`quantization_config` the exporter set on the config object, which
`save_config` alone would drop.

Signed-off-by: Nikita Davidchuk <ar4ikov228@gmail.com>
@changwangss
changwangss force-pushed the fix-diffusers-config-export branch from 4083419 to f6738c8 Compare September 16, 2026 06:32
@changwangss

Copy link
Copy Markdown
Contributor

@xin3he I rebased this PR onto the latest main, addressed the config-saving questions in the review threads, and reran the download-free tiny SD3 export regression (1 passed). Could you please take another look and dismiss the requested-changes review if the explanation and current implementation resolve your concerns?

@changwangss changwangss left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed after rebase onto current main. The direct Diffusers ModelMixin path is distinct from the pipeline loader patch, delegates to Diffusers model.save_config, preserves quantization_config, and passes the tiny SD3 plus immediate-saving regression tests.

@changwangss
changwangss requested a review from xin3he September 16, 2026 06:40
@AutoRoundBot

Copy link
Copy Markdown
Collaborator

/azp run Unit-Test-CUDA-AutoRound

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@changwangss

Copy link
Copy Markdown
Contributor

/azp run Unit-Test-CUDA-AutoRound

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@xin3he xin3he left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@xin3he
xin3he requested a lite review from Copilot September 16, 2026 07:51
@changwangss
changwangss merged commit 1983b65 into intel:main Sep 16, 2026
49 checks passed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Two moderate findings remain around quantization metadata being lost or overwritten in non-immediate save paths.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Fixes diffusers ModelMixin export config serialization and preserves quantization metadata.

Changes:

  • Adds shared diffusers-aware config-saving logic.
  • Updates fake-format and LLM Compressor export paths.
  • Adds an SD3 regression test.
File summaries
File Summary
test/unit/test_cpu/export/test_export.py Adds diffusers config export regression coverage.
auto_round/export/utils.py Adds shared config saving. Moderate finding (2 votes): non-immediate exports can still omit quantization metadata.
auto_round/export/formats/backends/fake.py Routes meta-device exports through the shared helper.
auto_round/export/export_to_llmcompressor/export.py Reuses the helper. Moderate finding (2 votes): a subsequent save can overwrite the merged metadata.
Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 2
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +382 to 384
save_config_artifact(model, output_dir)

save_model(model, output_dir, safe_serialization=safe_serialization, immediate_saving=immediate_saving)
Comment on lines +79 to +80
def _save_model_configs(model: nn.Module, save_dir: str) -> None:
save_config_artifact(model, save_dir)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants