Skip to content

[BUG] GPTQ-quantized OVIS 1B model yields poor performance & misaligned outputs in vLLM-0.9.1 #1653

Description

@AstonyJ

https://huggingface.co/AIDC-AI/Ovis2-2B-GPTQ-Int4

The OVIS 1B model quantized using the above GPTQ code performs extremely poorly when accelerated with vLLM-0.9.1, and the output precision is completely inconsistent. What could be the reason?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions