fix: save the config of diffusers models on export - #2119
Conversation
|
@changwangss could you help fix this issue if the author is not free? |
The export path ended with `model.config.save_pretrained(save_dir)`. A diffusers `ModelMixin` keeps its config in a `FrozenDict`, which has no such method, so `quantize_and_save` on one wrote every shard and then died with an AttributeError: no config.json, no quantization_config, nothing that can be loaded back. The diffusion path escapes this only because `diffusion_load_model` patches a `save_pretrained` onto the pipeline config, so anyone who loads a transformer themselves and hands it to AutoRound gets the half-written checkpoint. Route all three config-saving sites through one helper that uses `ModelMixin.save_config` for diffusers models and merges back the `quantization_config` the exporter set on the config object, which `save_config` alone would drop. Signed-off-by: Nikita Davidchuk <ar4ikov228@gmail.com>
4083419 to
f6738c8
Compare
|
@xin3he I rebased this PR onto the latest |
changwangss
left a comment
There was a problem hiding this comment.
Reviewed after rebase onto current main. The direct Diffusers ModelMixin path is distinct from the pipeline loader patch, delegates to Diffusers model.save_config, preserves quantization_config, and passes the tiny SD3 plus immediate-saving regression tests.
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
There was a problem hiding this comment.
🟡 Changes recommended
Two moderate findings remain around quantization metadata being lost or overwritten in non-immediate save paths.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Fixes diffusers ModelMixin export config serialization and preserves quantization metadata.
Changes:
- Adds shared diffusers-aware config-saving logic.
- Updates fake-format and LLM Compressor export paths.
- Adds an SD3 regression test.
File summaries
| File | Summary |
|---|---|
test/unit/test_cpu/export/test_export.py |
Adds diffusers config export regression coverage. |
auto_round/export/utils.py |
Adds shared config saving. Moderate finding (2 votes): non-immediate exports can still omit quantization metadata. |
auto_round/export/formats/backends/fake.py |
Routes meta-device exports through the shared helper. |
auto_round/export/export_to_llmcompressor/export.py |
Reuses the helper. Moderate finding (2 votes): a subsequent save can overwrite the merged metadata. |
Review details
- Files reviewed: 4/4 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| save_config_artifact(model, output_dir) | ||
|
|
||
| save_model(model, output_dir, safe_serialization=safe_serialization, immediate_saving=immediate_saving) |
| def _save_model_configs(model: nn.Module, save_dir: str) -> None: | ||
| save_config_artifact(model, save_dir) |
Description
quantize_and_saveon a diffusersModelMixinwrites every safetensors shard and then dies before writing any config, leaving a checkpoint nothing can load.A diffusers model's
.configis adiffusers.configuration_utils.FrozenDict, which has nosave_pretrained;ModelMixin.save_configis the equivalent writer. The diffusion path never hits this becausediffusion_load_modelpatches asave_pretrainedonto the pipeline and component configs (auto_round/utils/model.py, around theconfig_save_pretrainedpartials). So the crash is reserved for callers who load a transformer themselves and pass it toAutoRound, and what they get is shards on disk with noconfig.jsonand noquantization_config.One detail that made the obvious fix wrong: calling
model.save_config(save_dir)alone writes aconfig.jsonwithoutquantization_config, because the exporter sets that on the config object (export_to_autoround/export.py:model.config.quantization_config = quantization_config) andsave_configonly serializes the config's own dict. I verified that directly —save_configafter setting the attribute produces aconfig.jsonwhose keys do not includequantization_config. So the helper writes the config with diffusers' own writer and then mergesquantization_configback in.The same
model.config.save_pretrained(...)call exists at two more sites, so all three now go through one helper:formats.py(fake format, meta-device branch) andexport_to_llmcompressor/export.py. I only exercised theauto_roundformat path with a diffusers model; the other two are the same substitution, unverified against a diffusers model.export_to_mlxhas the same call and is left alone — a diffusers transformer is not an MLX target.Reproduction
Against
mainat 417d87d, with torch 2.11.0+cu128 / diffusers 0.40.0.dev0 / transformers 5.14.1:Before:
AttributeErroras above, and the only thing on disk isw4g32/model.safetensors.After: no error, and
with
config.jsoncarrying_class_name: SD3Transformer2DModel,_diffusers_version, and aquantization_config(quant_method: auto-round,bits: 4,group_size: 32,packing_format: auto_round:auto_gptq).Two honest notes about that snippet:
_StubTokenizeris not part of this bug. On currentmain,ModelContextraisesValueError("A tokenizer must be set for non-str model input")for any model object, including data-free RTN, becauseBaseOrchestrator.__init__passesneed_calib=self.need_calib(still the class defaultTrue) beforeself.need_calib = self._check_need_calib()runs later in the same__init__. That looks like a separate ordering bug; happy to file it separately.model-*.safetensors, while diffusers looks fordiffusion_pytorch_model*. This PR does not touch naming —rename_weights_filesalready exists for that and is used by the diffusion path.Test
test/test_cpu/export/test_export.py::test_save_model_writes_diffusers_configbuilds a tiny SD3 transformer in-process (no download) and callssave_model(..., immediate_saving=True), the path that crashed.Where this came from
Quantizing MiniMax-H3's 33B video/audio DiT (
MiniMaxH3Transformer3DModel) to W4A16. The transformer is loaded directly rather than through a pipeline, so the export died after writing all 14 shards; the config had to be reconstructed by hand from the tensors on disk. Result and write-up: https://huggingface.co/Ar4ikov/MiniMax-H3-transformer-W4A16-RTNType of Change
Bug fix
Related Issues
No issue filed — the failure is deterministic and the fix is small. Happy to open one if you prefer it tracked.
Checklist Before Submitting
autogptq_format,autoround_format,awq_format,immediate_saving) the patched tree gives the same result as unpatchedmainin the same environment:5 failed, 1 passedeither way, plus the new test passing. All five failures arecheck_compressed_tensors_supported(ImportError: Please install compressed-tensors .../SystemExit: -1) — a missing optional dependency here, which also means thellm_compressorexport path is not covered by my local run./azp run Unit-Test-CUDA-AutoRoundon this repo; please kick it off.@WeiweiZhang1, you last touched this config-saving code in #1810, and @mengniwang95 as the author of diffusion model saving (#1519), where the
save_pretrained-on-the-config patch that hides this bug comes from.