Skip to content

Questions about HPSv3 scorer correctness, Figure 2 sample quality, and DiffusionNFT baseline #7

Description

@Shi-wang-MIT

Hi, thanks for releasing the code. I have a few questions regarding the implementation and the reported qualitative results.

1. HPSv3 scorer implementation appears to miss image normalization

I think there may be a serious correctness issue in the current HPSv3 scorer.

In src/hpsv3_scorer.py, the differentiable preprocessing path calls:

self.ip._preprocess(images01[i:i + 1], do_rescale=False)

However, in the official HPSv3 differentiable image processor, _preprocess() does not resolve do_normalize=None, image_mean=None, and image_std=None to the processor defaults. That resolution only happens in the public preprocess() / preprocess_tensor() path.

As a result, although disabling rescaling is appropriate for an input already in [0, 1], the image does not appear to receive the required CLIP mean/std normalization before being fed into the Qwen2-VL vision encoder.

This seems potentially quite significant: the resulting HPSv3 scores, and especially the image-space reward gradients used for optimization, may not correspond to the official HPSv3 scorer.

Have you checked numerical parity between this implementation and the official

HPSv3RewardInferencer.reward(...)

on exactly the same image/prompt pairs?

It would be helpful if you could provide a simple parity test comparing the raw HPSv3 scores from the released scorer against the official implementation.

2. Figure 2 / teaser image quality

I also have a question about the qualitative results in the paper.

The samples shown in Figure 2 / the main qualitative figure appear substantially higher quality than what I obtain from the released implementation and than some of the other reported qualitative results.

Could you clarify exactly how these images were generated?

In particular, were they generated using exactly the same released checkpoint and inference configuration? It would be useful to provide the corresponding:

  • prompts,
  • random seeds,
  • checkpoints,
  • sampling steps,
  • CFG/guidance settings,
  • resolution, and
  • any sample-selection or curation procedure.

This would make the qualitative comparison much easier to reproduce.

3. DiffusionNFT baseline outputs are consistently blurry

Finally, I am having difficulty reproducing a reasonable DiffusionNFT baseline using the released implementation.

The images generated by the provided DiffusionNFT baseline are consistently very blurry / low quality in my runs. This seems unusual enough that I am concerned there may be an implementation or inference-configuration issue with the baseline.

Could you clarify whether you verified this implementation against the original DiffusionNFT implementation?

In particular, could you provide the exact DiffusionNFT training and inference configuration used for the paper, as well as some representative baseline generations? It would also be useful to confirm that DiffusionNFT and DiffusionOPSD are evaluated using the same sampling resolution, number of steps, guidance settings, and other inference hyperparameters.

Thanks — I would appreciate any clarification on these points, especially the HPSv3 preprocessing issue, since that may affect both the reported HPSv3 evaluation numbers and optimization results.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions