Super-squash branch 'main' using huggingface_hub

Browse files

Files changed (15) hide show

.gitattributes +36 -0
README.md +204 -0
chat_template.jinja +140 -0
config.json +276 -0
generation_config.json +15 -0
model-00001-of-00003.safetensors +3 -0
model-00002-of-00003.safetensors +3 -0
model-00003-of-00003.safetensors +3 -0
model.safetensors.index.json +0 -0
preprocessor_config.json +11 -0
recipe.yaml +37 -0
special_tokens_map.json +42 -0
tokenizer.json +3 -0
tokenizer_config.json +327 -0
video_preprocessor_config.json +11 -0

.gitattributes ADDED Viewed

	@@ -0,0 +1,36 @@

+*.7z filter=lfs diff=lfs merge=lfs -text
+*.arrow filter=lfs diff=lfs merge=lfs -text
+*.bin filter=lfs diff=lfs merge=lfs -text
+*.bz2 filter=lfs diff=lfs merge=lfs -text
+*.ckpt filter=lfs diff=lfs merge=lfs -text
+*.ftz filter=lfs diff=lfs merge=lfs -text
+*.gz filter=lfs diff=lfs merge=lfs -text
+*.h5 filter=lfs diff=lfs merge=lfs -text
+*.joblib filter=lfs diff=lfs merge=lfs -text
+*.lfs.* filter=lfs diff=lfs merge=lfs -text
+*.mlmodel filter=lfs diff=lfs merge=lfs -text
+*.model filter=lfs diff=lfs merge=lfs -text
+*.msgpack filter=lfs diff=lfs merge=lfs -text
+*.npy filter=lfs diff=lfs merge=lfs -text
+*.npz filter=lfs diff=lfs merge=lfs -text
+*.onnx filter=lfs diff=lfs merge=lfs -text
+*.ot filter=lfs diff=lfs merge=lfs -text
+*.parquet filter=lfs diff=lfs merge=lfs -text
+*.pb filter=lfs diff=lfs merge=lfs -text
+*.pickle filter=lfs diff=lfs merge=lfs -text
+*.pkl filter=lfs diff=lfs merge=lfs -text
+*.pt filter=lfs diff=lfs merge=lfs -text
+*.pth filter=lfs diff=lfs merge=lfs -text
+*.rar filter=lfs diff=lfs merge=lfs -text
+*.safetensors filter=lfs diff=lfs merge=lfs -text
+saved_model/**/* filter=lfs diff=lfs merge=lfs -text
+*.tar.* filter=lfs diff=lfs merge=lfs -text
+*.tar filter=lfs diff=lfs merge=lfs -text
+*.tflite filter=lfs diff=lfs merge=lfs -text
+*.tgz filter=lfs diff=lfs merge=lfs -text
+*.wasm filter=lfs diff=lfs merge=lfs -text
+*.xz filter=lfs diff=lfs merge=lfs -text
+*.zip filter=lfs diff=lfs merge=lfs -text
+*.zst filter=lfs diff=lfs merge=lfs -text
+*tfevents* filter=lfs diff=lfs merge=lfs -text
+tokenizer.json filter=lfs diff=lfs merge=lfs -text

README.md ADDED Viewed

	@@ -0,0 +1,204 @@

+---
+language:
+- zh
+- en
+library_name: transformers
+license: mit
+pipeline_tag: image-text-to-text
+base_model: zai-org/GLM-4.6V-Flash
+---
+# GLM-4.6V-Flash AWQ - INT8
+## Model Details
+### Quantization Details
+- **Quantization Method:** AWQ
+- **Bits:** 8
+- **Group Size:** 32
+- **Calibration Dataset:** [5CD-AI/LLaVA-CoT-o1-Instruct](https://huggingface.co/datasets/5CD-AI/LLaVA-CoT-o1-Instruct)
+- **Quantization Tool:** [llm-compressor](https://github.com/vllm-project/llm-compressor)
+### Memory Usage
+| **Type** | **GLM-4.6V-Flash** | **GLM-4.6V-Flash-AWQ-8bit** |
+|:---------------:|:----------------:|:----------------:|
+| **Memory Size** | 19.2 GB | 12.0 GB |
+| **KV Cache per Token** | 40.0 kB | 20.0 kB |
+| **KV Cache per Context** | 5.0 GB | 2.5 GB |
+## Inference
+### Prerequisite
+```bash
+pip install vllm>=0.12.0
+pip install --upgrade git+https://github.com/huggingface/transformers.git
+```
+### Basic Usage
+```bash
+vllm serve cyankiwi/GLM-4.6V-Flash-AWQ-8bit
+```
+## Additional Information
+### Changelog
+- **v1.0.0** - Initial quantized release
+### Authors
+- **Name:** Ton Cao
+- **Contacts:** [email protected]
+# GLM-4.6V
+<div align="center">
+<img src=https://raw.githubusercontent.com/zai-org/GLM-V/refs/heads/main/resources/logo.svg width="40%"/>
+</div>
+This model is part of the GLM-V family of models, introduced in the paper [GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning](https://huggingface.co/papers/2507.01006).
+-   **GLM-4.6V Blog**: [https://z.ai/blog/glm-4.6v](https://z.ai/blog/glm-4.6v)
+-   **Paper**: [https://huggingface.co/papers/2507.01006](https://huggingface.co/papers/2507.01006)
+-   **GitHub Repository**: [https://github.com/zai-org/GLM-V](https://github.com/zai-org/GLM-V)
+-   **Online Demo**: [https://chat.z.ai/](https://chat.z.ai/)
+-   **API Access**: [Z.ai Open Platform](https://docs.z.ai/guides/vlm/glm-4.6v)
+-   **Desktop Assistant App**: [https://huggingface.co/spaces/zai-org/GLM-4.5V-Demo-App](https://huggingface.co/spaces/zai-org/GLM-4.5V-Demo-App)
+## Introduction
+GLM-4.6V series model includes two versions: GLM-4.6V (106B), a foundation model designed for cloud and high-performance
+cluster scenarios,
+and GLM-4.6V-Flash (9B), a lightweight model optimized for local deployment and low-latency applications.
+GLM-4.6V scales its context window to 128k tokens in training,
+and achieves SoTA performance in visual understanding among models of similar parameter scales.
+Crucially, we integrate native Function Calling capabilities for the first time.
+This effectively bridges the gap between "visual perception" and "executable action"
+providing a unified technical foundation for multimodal agents in real-world business scenarios.
+![GLM-4.6V Benchmarks](https://raw.githubusercontent.com/zai-org/GLM-V/refs/heads/main/resources/bench_46v.jpeg)
+Beyond achieves SoTA performance across major multimodal benchmarks at comparable model scales. GLM-4.6V introduces
+several key features:
+- **Native Multimodal Function Calling**
+Enables native vision-driven tool use. Images, screenshots, and document pages can be passed directly as tool inputs without text conversion, while visual outputs (charts, search images, rendered pages) are interpreted and integrated into the reasoning chain. This closes the loop from perception to understanding to execution.
+- **Interleaved Image-Text Content Generation**
+Supports high-quality mixed media creation from complex multimodal inputs. GLM-4.6V takes a multimodal context—spanning documents, user inputs, and tool-retrieved images—and synthesizes coherent, interleaved image-text content tailored to the task. During generation it can actively call search and retrieval tools to gather and curate additional text and visuals, producing rich, visually grounded content.
+- **Multimodal Document Understanding**
+GLM-4.6V can process up to 128K tokens of multi-document or long-document input, directly interpreting richly formatted pages as images. It understands text, layout, charts, tables, and figures jointly, enabling accurate comprehension of complex, image-heavy documents without requiring prior conversion to plain text.
+- **Frontend Replication & Visual Editing**
+Reconstructs pixel-accurate HTML/CSS from UI screenshots and supports natural-language-driven edits. It detects layout, components, and styles visually, generates clean code, and applies iterative visual modifications through simple user instructions.
+**This Hugging Face repository hosts the `GLM-4.6V-Flash` model, part of the `GLM-V` series.**
+## Usage
+### Environment Installation
+For `SGLang`:
+```bash
+pip install sglang>=0.5.6post1
+pip install transformers>=5.0.0rc0
+```
+For `vLLM`:
+```bash
+pip install vllm>=0.12.0
+pip install transformers>=5.0.0rc0
+```
+### Quick Start with Transformers
+```python
+from transformers import AutoProcessor, Glm4vMoeForConditionalGeneration
+import torch
+MODEL_PATH = "zai-org/GLM-4.6V-Flash"
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {
+                "type": "image",
+                "url": "https://upload.wikimedia.org/wikipedia/commons/f/fa/Grayscale_8bits_palette_sample_image.png"
+            },
+            {
+                "type": "text",
+                "text": "describe this image"
+            }
+        ],
+    }
+]
+processor = AutoProcessor.from_pretrained(MODEL_PATH)
+model = Glm4vMoeForConditionalGeneration.from_pretrained(
+    pretrained_model_name_or_path=MODEL_PATH,
+    torch_dtype="auto",
+    device_map="auto",
+)
+inputs = processor.apply_chat_template(
+    messages,
+    tokenize=True,
+    add_generation_prompt=True,
+    return_dict=True,
+    return_tensors="pt"
+).to(model.device)
+inputs.pop("token_type_ids", None)
+generated_ids = model.generate(**inputs, max_new_tokens=8192)
+output_text = processor.decode(generated_ids[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)
+print(output_text)
+```
+## Evaluation Settings
+We primarily use vLLM as the backend for model inference. For faster and more reliable performance on video tasks, we employ SGLang. To reproduce our leaderboard results, we recommend the following decoding parameters:
++	top_p: 0.6
++	top_k: 2
++	temperature: 0.8
++	repetition_penalty: 1.1
++	max_generate_tokens: 16K
+For more usage details, please refer to Our [Github](https://github.com/zai-org/GLM-V).
+## Fixed and Remaining Issues
+Since the open-sourcing of GLM-4.1V, we have received extensive feedback from the community and are well aware that the model still has many shortcomings. In subsequent iterations, we attempted to address several common issues — such as repetitive thinking outputs and formatting errors — which have been mitigated to some extent in this new version.
+However, the model still has several limitations and issues that we will fix as soon as possible:
+1. Pure text QA capabilities still have significant room for improvement. In this development cycle, our primary focus was on visual multimodal scenarios, and we will enhance pure text abilities in upcoming updates.
+2. The model may still overthink or even repeat itself in certain cases, especially when dealing with complex prompts.
+3. In some situations, the model may restate the answer again at the end.
+4. There remain certain perception limitations, such as counting accuracy and identifying specific individuals, which still require improvement.
+Thank you for your patience and understanding. We also welcome feedback and suggestions in the issue section — we will respond and improve as much as we can!
+## Citation
+If you use this model, please cite the following paper:
+```bibtex
+@misc{vteam2025glm45vglm41vthinkingversatilemultimodal,
+      title={GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning},
+      author={V Team and Wenyi Hong and Wenmeng Yu and Xiaotao Gu and Guo Wang and Guobing Gan and Haomiao Tang and Jiale Cheng and Ji Qi and Junhui Ji and Lihang Pan and Shuaiqi Duan and Weihan Wang and Yan Wang and Yean Cheng and Zehai He and Zhe Su and Zhen Yang and Ziyang Pan and Aohan Zeng and Baoxu Wang and Bin Chen and Boyan Shi and Changyu Pang and Chenhui Zhang and Da Yin and Fan Yang and Guoqing Chen and Jiazheng Xu and Jiale Zhu and Jiali Chen and Jing Chen and Jinhao Chen and Jinghao Lin and Jinjiang Wang and Junjie Chen and Leqi Lei and Letian Gong and Leyi Pan and Mingdao Liu and Mingde Xu and Mingzhi Zhang and Qinkai Zheng and Sheng Yang and Shi Zhong and Shiyu Huang and Shuyuan Zhao and Siyan Xue and Shangqin Tu and Shengbiao Meng and Tianshu Zhang and Tianwei Luo and Tianxiang Hao and Tianyu Tong and Wenkai Li and Wei Jia and Xiao Liu and Xiaohan Zhang and Xin Lyu and Xinyue Fan and Xuancheng Huang and Yanling Wang and Yadong Xue and Yanfeng Wang and Yanzi Wang and Yifan An and Yifan Du and Yiming Shi and Yiheng Huang and Yilin Niu and Yuan Wang and Yuanchang Yue and Yuchen Li and Yutao Zhang and Yuting Wang and Yu Wang and Yuxuan Zhang and Zhao Xue and Zhenyu Hou and Zhengxiao Du and Zihan Wang and Peng Zhang and Debing Liu and Bin Xu and Juanzi Li and Minlie Huang and Yuxiao Dong and Jie Tang},
+      year={2025},
+      eprint={2507.01006},
+      archivePrefix={arXiv},
+      primaryClass={cs.CV},
+      url={https://arxiv.org/abs/2507.01006},
+}
+```

chat_template.jinja ADDED Viewed

	@@ -0,0 +1,140 @@

+[gMASK]<sop>
+{%- if tools -%}
+<|system|>
+# Tools
+You may call one or more functions to assist with the user query.
+You are provided with function signatures within <tools></tools> XML tags:
+<tools>
+{% for tool in tools %}
+{{ tool | tojson(ensure_ascii=False) }}
+{% endfor %}
+</tools>
+For each function call, output the function name and arguments within the following XML format:
+<tool_call>{function-name}
+<arg_key>{arg-key-1}</arg_key>
+<arg_value>{arg-value-1}</arg_value>
+<arg_key>{arg-key-2}</arg_key>
+<arg_value>{arg-value-2}</arg_value>
+...
+</tool_call>{%- endif -%}
+{%- macro visible_text(content) -%}
+    {%- if content is string -%}
+        {{- content }}
+    {%- elif content is iterable and content is not mapping -%}
+        {%- for item in content -%}
+            {%- if item is mapping and item.type == 'text' -%}
+                {{- item.text }}
+            {%- elif item is mapping and (item.type == 'image' or 'image' in item) -%}
+                <|begin_of_image|><|image|><|end_of_image|>
+            {%- elif item is mapping and (item.type == 'video' or 'video' in item) -%}
+                <|begin_of_video|><|video|><|end_of_video|>
+            {%- elif item is string -%}
+                {{- item }}
+            {%- endif -%}
+        {%- endfor -%}
+    {%- else -%}
+        {{- content }}
+    {%- endif -%}
+{%- endmacro -%}
+{%- set ns = namespace(last_user_index=-1) %}
+{%- for m in messages %}
+    {%- if m.role == 'user' %}
+        {% set ns.last_user_index = loop.index0 -%}
+    {%- endif %}
+{%- endfor %}
+{% for m in messages %}
+{%- if m.role == 'user' -%}<|user|>
+{% if m.content is string %}
+{{ m.content }}
+{%- else %}
+{%- for item in m.content %}
+{% if item.type == 'video' or 'video' in item %}
+<|begin_of_video|><|video|><|end_of_video|>{% elif item.type == 'image' or 'image' in item %}
+<|begin_of_image|><|image|><|end_of_image|>{% elif item.type == 'text' %}
+{{ item.text }}
+{%- endif %}
+{%- endfor %}
+{%- endif %}
+{{- '/nothink' if (enable_thinking is defined and not enable_thinking and not visible_text(m.content).endswith("/nothink")) else '' -}}
+{%- elif m.role == 'assistant' -%}
+<|assistant|>
+{%- set reasoning_content = '' %}
+{%- set content = visible_text(m.content) %}
+{%- if m.reasoning_content is string %}
+    {%- set reasoning_content = m.reasoning_content %}
+{%- else %}
+    {%- if '</think>' in content %}
+        {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
+        {%- set content = content.split('</think>')[-1].lstrip('\n') %}
+    {%- endif %}
+{%- endif %}
+{%- if loop.index0 > ns.last_user_index and reasoning_content -%}
+{{ '\n<think>' + reasoning_content.strip() +  '</think>'}}
+{%- else -%}
+{{ '\n<think></think>' }}
+{%- endif -%}
+{%- if content.strip() -%}
+{{ '\n' + content.strip() }}
+{%- endif -%}
+{% if m.tool_calls %}
+{% for tc in m.tool_calls %}
+{%- if tc.function %}
+    {%- set tc = tc.function %}
+{%- endif %}
+{{ '\n<tool_call>' + tc.name }}
+{% set _args = tc.arguments %}
+{% for k, v in _args.items() %}
+<arg_key>{{ k }}</arg_key>
+<arg_value>{{ v | tojson(ensure_ascii=False) if v is not string else v }}</arg_value>
+{% endfor %}
+</tool_call>{% endfor %}
+{% endif %}
+{%- elif m.role == 'tool' -%}
+{%- if m.content is string -%}
+{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
+    {{- '<|observation|>' }}
+{%- endif %}
+{{- '\n<tool_response>\n' }}
+{{- m.content }}
+{{- '\n</tool_response>' }}
+{% elif m.content is iterable and m.content is not mapping %}
+{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
+{{- '<|observation|>' }}
+{%- endif %}
+{{- '\n<tool_response>\n' }}
+{%- for tr in m.content -%}
+  {%- if tr is mapping and tr.type is defined -%}
+    {%- set t = tr.type | lower -%}
+    {%- if t == 'text' and tr.text is defined -%}
+{{ tr.text }}
+    {%- elif t in ['image', 'image_url'] -%}
+<|begin_of_image|><|image|><|end_of_image|>
+    {%- elif t in ['video', 'video_url'] -%}
+<|begin_of_video|><|video|><|end_of_video|>
+    {%- else -%}
+{{ tr | tojson(ensure_ascii=False) }}
+    {%- endif -%}
+  {%- else -%}
+{{ tr.output if tr.output is defined else tr }}
+  {%- endif -%}
+{%- endfor -%}
+{{- '\n</tool_response>' }}
+{%- else -%}
+<|observation|>{% for tr in m.content %}
+<tool_response>
+{{ tr.output if tr.output is defined else tr }}
+</tool_response>{% endfor -%}
+{% endif -%}
+{%- elif m.role == 'system' -%}
+<|system|>
+{{ visible_text(m.content) }}
+{%- endif -%}
+{%- endfor -%}
+{%- if add_generation_prompt -%}
+<|assistant|>
+{{'<think></think>\n' if (enable_thinking is defined and not enable_thinking) else ''}}
+{%- endif -%}

config.json ADDED Viewed

	@@ -0,0 +1,276 @@

+{
+  "architectures": [
+    "Glm4vForConditionalGeneration"
+  ],
+  "dtype": "float16",
+  "image_end_token_id": 151340,
+  "image_start_token_id": 151339,
+  "image_token_id": 151363,
+  "model_type": "glm4v",
+  "quantization_config": {
+    "config_groups": {
+      "group_0": {
+        "format": "pack-quantized",
+        "input_activations": null,
+        "output_activations": null,
+        "targets": [
+          "Linear"
+        ],
+        "weights": {
+          "actorder": null,
+          "block_structure": null,
+          "dynamic": false,
+          "group_size": 32,
+          "num_bits": 8,
+          "observer": "mse",
+          "observer_kwargs": {},
+          "strategy": "group",
+          "symmetric": true,
+          "type": "int"
+        }
+      }
+    },
+    "format": "pack-quantized",
+    "global_compression_ratio": null,
+    "ignore": [
+      "model.visual.blocks.0.attn.qkv_proj",
+      "model.visual.blocks.0.attn.qkv",
+      "model.visual.blocks.0.attn.proj",
+      "model.visual.blocks.0.mlp.gate_up_proj",
+      "model.visual.blocks.0.mlp.gate_proj",
+      "model.visual.blocks.0.mlp.up_proj",
+      "model.visual.blocks.0.mlp.down_proj",
+      "model.visual.blocks.1.attn.qkv_proj",
+      "model.visual.blocks.1.attn.qkv",
+      "model.visual.blocks.1.attn.proj",
+      "model.visual.blocks.1.mlp.gate_up_proj",
+      "model.visual.blocks.1.mlp.gate_proj",
+      "model.visual.blocks.1.mlp.up_proj",
+      "model.visual.blocks.1.mlp.down_proj",
+      "model.visual.blocks.2.attn.qkv_proj",
+      "model.visual.blocks.2.attn.qkv",
+      "model.visual.blocks.2.attn.proj",
+      "model.visual.blocks.2.mlp.gate_up_proj",
+      "model.visual.blocks.2.mlp.gate_proj",
+      "model.visual.blocks.2.mlp.up_proj",
+      "model.visual.blocks.2.mlp.down_proj",
+      "model.visual.blocks.3.attn.qkv_proj",
+      "model.visual.blocks.3.attn.qkv",
+      "model.visual.blocks.3.attn.proj",
+      "model.visual.blocks.3.mlp.gate_up_proj",
+      "model.visual.blocks.3.mlp.gate_proj",
+      "model.visual.blocks.3.mlp.up_proj",
+      "model.visual.blocks.3.mlp.down_proj",
+      "model.visual.blocks.4.attn.qkv_proj",
+      "model.visual.blocks.4.attn.qkv",
+      "model.visual.blocks.4.attn.proj",
+      "model.visual.blocks.4.mlp.gate_up_proj",
+      "model.visual.blocks.4.mlp.gate_proj",
+      "model.visual.blocks.4.mlp.up_proj",
+      "model.visual.blocks.4.mlp.down_proj",
+      "model.visual.blocks.5.attn.qkv_proj",
+      "model.visual.blocks.5.attn.qkv",
+      "model.visual.blocks.5.attn.proj",
+      "model.visual.blocks.5.mlp.gate_up_proj",
+      "model.visual.blocks.5.mlp.gate_proj",
+      "model.visual.blocks.5.mlp.up_proj",
+      "model.visual.blocks.5.mlp.down_proj",
+      "model.visual.blocks.6.attn.qkv_proj",
+      "model.visual.blocks.6.attn.qkv",
+      "model.visual.blocks.6.attn.proj",
+      "model.visual.blocks.6.mlp.gate_up_proj",
+      "model.visual.blocks.6.mlp.gate_proj",
+      "model.visual.blocks.6.mlp.up_proj",
+      "model.visual.blocks.6.mlp.down_proj",
+      "model.visual.blocks.7.attn.qkv_proj",
+      "model.visual.blocks.7.attn.qkv",
+      "model.visual.blocks.7.attn.proj",
+      "model.visual.blocks.7.mlp.gate_up_proj",
+      "model.visual.blocks.7.mlp.gate_proj",
+      "model.visual.blocks.7.mlp.up_proj",
+      "model.visual.blocks.7.mlp.down_proj",
+      "model.visual.blocks.8.attn.qkv_proj",
+      "model.visual.blocks.8.attn.qkv",
+      "model.visual.blocks.8.attn.proj",
+      "model.visual.blocks.8.mlp.gate_up_proj",
+      "model.visual.blocks.8.mlp.gate_proj",
+      "model.visual.blocks.8.mlp.up_proj",
+      "model.visual.blocks.8.mlp.down_proj",
+      "model.visual.blocks.9.attn.qkv_proj",
+      "model.visual.blocks.9.attn.qkv",
+      "model.visual.blocks.9.attn.proj",
+      "model.visual.blocks.9.mlp.gate_up_proj",
+      "model.visual.blocks.9.mlp.gate_proj",
+      "model.visual.blocks.9.mlp.up_proj",
+      "model.visual.blocks.9.mlp.down_proj",
+      "model.visual.blocks.10.attn.qkv_proj",
+      "model.visual.blocks.10.attn.qkv",
+      "model.visual.blocks.10.attn.proj",
+      "model.visual.blocks.10.mlp.gate_up_proj",
+      "model.visual.blocks.10.mlp.gate_proj",
+      "model.visual.blocks.10.mlp.up_proj",
+      "model.visual.blocks.10.mlp.down_proj",
+      "model.visual.blocks.11.attn.qkv_proj",
+      "model.visual.blocks.11.attn.qkv",
+      "model.visual.blocks.11.attn.proj",
+      "model.visual.blocks.11.mlp.gate_up_proj",
+      "model.visual.blocks.11.mlp.gate_proj",
+      "model.visual.blocks.11.mlp.up_proj",
+      "model.visual.blocks.11.mlp.down_proj",
+      "model.visual.blocks.12.attn.qkv_proj",
+      "model.visual.blocks.12.attn.qkv",
+      "model.visual.blocks.12.attn.proj",
+      "model.visual.blocks.12.mlp.gate_up_proj",
+      "model.visual.blocks.12.mlp.gate_proj",
+      "model.visual.blocks.12.mlp.up_proj",
+      "model.visual.blocks.12.mlp.down_proj",
+      "model.visual.blocks.13.attn.qkv_proj",
+      "model.visual.blocks.13.attn.qkv",
+      "model.visual.blocks.13.attn.proj",
+      "model.visual.blocks.13.mlp.gate_up_proj",
+      "model.visual.blocks.13.mlp.gate_proj",
+      "model.visual.blocks.13.mlp.up_proj",
+      "model.visual.blocks.13.mlp.down_proj",
+      "model.visual.blocks.14.attn.qkv_proj",
+      "model.visual.blocks.14.attn.qkv",
+      "model.visual.blocks.14.attn.proj",
+      "model.visual.blocks.14.mlp.gate_up_proj",
+      "model.visual.blocks.14.mlp.gate_proj",
+      "model.visual.blocks.14.mlp.up_proj",
+      "model.visual.blocks.14.mlp.down_proj",
+      "model.visual.blocks.15.attn.qkv_proj",
+      "model.visual.blocks.15.attn.qkv",
+      "model.visual.blocks.15.attn.proj",
+      "model.visual.blocks.15.mlp.gate_up_proj",
+      "model.visual.blocks.15.mlp.gate_proj",
+      "model.visual.blocks.15.mlp.up_proj",
+      "model.visual.blocks.15.mlp.down_proj",
+      "model.visual.blocks.16.attn.qkv_proj",
+      "model.visual.blocks.16.attn.qkv",
+      "model.visual.blocks.16.attn.proj",
+      "model.visual.blocks.16.mlp.gate_up_proj",
+      "model.visual.blocks.16.mlp.gate_proj",
+      "model.visual.blocks.16.mlp.up_proj",
+      "model.visual.blocks.16.mlp.down_proj",
+      "model.visual.blocks.17.attn.qkv_proj",
+      "model.visual.blocks.17.attn.qkv",
+      "model.visual.blocks.17.attn.proj",
+      "model.visual.blocks.17.mlp.gate_up_proj",
+      "model.visual.blocks.17.mlp.gate_proj",
+      "model.visual.blocks.17.mlp.up_proj",
+      "model.visual.blocks.17.mlp.down_proj",
+      "model.visual.blocks.18.attn.qkv_proj",
+      "model.visual.blocks.18.attn.qkv",
+      "model.visual.blocks.18.attn.proj",
+      "model.visual.blocks.18.mlp.gate_up_proj",
+      "model.visual.blocks.18.mlp.gate_proj",
+      "model.visual.blocks.18.mlp.up_proj",
+      "model.visual.blocks.18.mlp.down_proj",
+      "model.visual.blocks.19.attn.qkv_proj",
+      "model.visual.blocks.19.attn.qkv",
+      "model.visual.blocks.19.attn.proj",
+      "model.visual.blocks.19.mlp.gate_up_proj",
+      "model.visual.blocks.19.mlp.gate_proj",
+      "model.visual.blocks.19.mlp.up_proj",
+      "model.visual.blocks.19.mlp.down_proj",
+      "model.visual.blocks.20.attn.qkv_proj",
+      "model.visual.blocks.20.attn.qkv",
+      "model.visual.blocks.20.attn.proj",
+      "model.visual.blocks.20.mlp.gate_up_proj",
+      "model.visual.blocks.20.mlp.gate_proj",
+      "model.visual.blocks.20.mlp.up_proj",
+      "model.visual.blocks.20.mlp.down_proj",
+      "model.visual.blocks.21.attn.qkv_proj",
+      "model.visual.blocks.21.attn.qkv",
+      "model.visual.blocks.21.attn.proj",
+      "model.visual.blocks.21.mlp.gate_up_proj",
+      "model.visual.blocks.21.mlp.gate_proj",
+      "model.visual.blocks.21.mlp.up_proj",
+      "model.visual.blocks.21.mlp.down_proj",
+      "model.visual.blocks.22.attn.qkv_proj",
+      "model.visual.blocks.22.attn.qkv",
+      "model.visual.blocks.22.attn.proj",
+      "model.visual.blocks.22.mlp.gate_up_proj",
+      "model.visual.blocks.22.mlp.gate_proj",
+      "model.visual.blocks.22.mlp.up_proj",
+      "model.visual.blocks.22.mlp.down_proj",
+      "model.visual.blocks.23.attn.qkv_proj",
+      "model.visual.blocks.23.attn.qkv",
+      "model.visual.blocks.23.attn.proj",
+      "model.visual.blocks.23.mlp.gate_up_proj",
+      "model.visual.blocks.23.mlp.gate_proj",
+      "model.visual.blocks.23.mlp.up_proj",
+      "model.visual.blocks.23.mlp.down_proj",
+      "model.visual.merger.proj",
+      "model.visual.merger.gate_up_proj",
+      "model.visual.merger.gate_proj",
+      "model.visual.merger.up_proj",
+      "model.visual.merger.down_proj",
+      "lm_head"
+    ],
+    "kv_cache_scheme": null,
+    "quant_method": "compressed-tensors",
+    "quantization_status": "compressed",
+    "sparsity_config": {},
+    "transform_config": {},
+    "version": "0.12.3.a20251203"
+  },
+  "text_config": {
+    "attention_bias": true,
+    "attention_dropout": 0.0,
+    "dtype": "bfloat16",
+    "eos_token_id": [
+      151329,
+      151336,
+      151338
+    ],
+    "hidden_act": "silu",
+    "hidden_size": 4096,
+    "image_token_id": null,
+    "initializer_range": 0.02,
+    "intermediate_size": 13696,
+    "max_position_embeddings": 131072,
+    "model_type": "glm4v_text",
+    "num_attention_heads": 32,
+    "num_hidden_layers": 40,
+    "num_key_value_heads": 2,
+    "pad_token_id": 151329,
+    "rms_norm_eps": 1e-05,
+    "rope_parameters": {
+      "mrope_section": [
+        8,
+        12,
+        12
+      ],
+      "partial_rotary_factor": 0.5,
+      "rope_theta": 500000,
+      "rope_type": "default"
+    },
+    "use_cache": true,
+    "vocab_size": 151552
+  },
+  "tie_word_embeddings": false,
+  "transformers_version": "4.57.3",
+  "video_end_token_id": 151342,
+  "video_start_token_id": 151341,
+  "video_token_id": 151364,
+  "vision_config": {
+    "attention_bias": false,
+    "attention_dropout": 0.0,
+    "depth": 24,
+    "hidden_act": "silu",
+    "hidden_dropout_prob": 0.0,
+    "hidden_size": 1536,
+    "image_size": 336,
+    "in_channels": 3,
+    "initializer_range": 0.02,
+    "intermediate_size": 13696,
+    "model_type": "glm4v",
+    "num_heads": 12,
+    "out_hidden_size": 4096,
+    "patch_size": 14,
+    "rms_norm_eps": 1e-05,
+    "spatial_merge_size": 2,
+    "temporal_patch_size": 2
+  }
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,15 @@

+{
+  "_from_model_config": true,
+  "do_sample": true,
+  "eos_token_id": [
+    151329,
+    151336,
+    151338,
+    151348
+  ],
+  "pad_token_id": 151329,
+  "temperature": 0.8,
+  "top_k": 2,
+  "top_p": 0.6,
+  "transformers_version": "4.57.3"
+}

model-00001-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:3ae8d7797802c40d5bcf90a92e0de138034a1f648f0a653d6cf7242a28ad243a
+size 4997257728

model-00002-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c410e1dbcc49ef10d2677620a41fe11ffb2e25bf34aa1cd3575b36ede05ac68f
+size 4985022776

model-00003-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:3b7696d696514be0224b306e1d42f97b7318259644fb92b19e6d787cec2fabe2
+size 2955378312

model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff

preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,11 @@

+{
+    "size": {"shortest_edge": 12544, "longest_edge": 9633792},
+    "do_rescale": true,
+    "patch_size": 14,
+    "temporal_patch_size": 2,
+    "merge_size": 2,
+    "image_mean": [0.48145466, 0.4578275, 0.40821073],
+    "image_std": [0.26862954, 0.26130258, 0.27577711],
+    "image_processor_type": "Glm46VImageProcessor",
+    "processor_class": "Glm46VProcessor"
+}

recipe.yaml ADDED Viewed

	@@ -0,0 +1,37 @@

+default_stage:
+  default_modifiers:
+    AWQModifier:
+      config_groups:
+        group_0:
+          targets: [Linear]
+          weights:
+            num_bits: 8
+            type: int
+            symmetric: true
+            group_size: 32
+            strategy: group
+            block_structure: null
+            dynamic: false
+            actorder: null
+            scale_dtype: null
+            zp_dtype: null
+            observer: mse
+            observer_kwargs: {}
+          input_activations: null
+          output_activations: null
+          format: null
+      targets: [Linear]
+      ignore: [lm_head, 're:.*embed_tokens', 're:.*input_layernorm', 're:.*post_attention_layernorm',
+        model.language_model.norm, 're:.*mlp[.]gate$', 're:model[.]visual.*']
+      mappings:
+      - smooth_layer: re:.*input_layernorm$
+        balance_layers: ['re:.*q_proj$', 're:.*k_proj$', 're:.*v_proj$']
+      - smooth_layer: re:.*v_proj$
+        balance_layers: ['re:.*o_proj$']
+      - smooth_layer: re:.*post_attention_layernorm$
+        balance_layers: ['re:.*gate_proj$', 're:.*up_proj$']
+      - smooth_layer: re:.*up_proj$
+        balance_layers: ['re:.*down_proj$']
+      offload_device: !!python/object/apply:torch.device [cpu]
+      duo_scaling: true
+      n_grid: 20

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,42 @@

+{
+  "additional_special_tokens": [
+    "<|endoftext|>",
+    "[MASK]",
+    "[gMASK]",
+    "[sMASK]",
+    "<sop>",
+    "<eop>",
+    "<|system|>",
+    "<|user|>",
+    "<|assistant|>",
+    "<|observation|>",
+    "<|begin_of_image|>",
+    "<|end_of_image|>",
+    "<|begin_of_video|>",
+    "<|end_of_video|>",
+    "<|begin_of_audio|>",
+    "<|end_of_audio|>",
+    "<|image|>",
+    "<|video|>",
+    "<|begin_of_transcription|>",
+    "<|end_of_transcription|>",
+    "<|code_prefix|>",
+    "<|code_middle|>",
+    "<|code_suffix|>",
+    "/nothink"
+  ],
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f2ff52959093921034528ecd6a59926e5fd543f56f94f2a0034ed4ba458c0a86
+size 19970698

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,327 @@

+{
+  "added_tokens_decoder": {
+    "151329": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151330": {
+      "content": "[MASK]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151331": {
+      "content": "[gMASK]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151332": {
+      "content": "[sMASK]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151333": {
+      "content": "<sop>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151334": {
+      "content": "<eop>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151335": {
+      "content": "<|system|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151336": {
+      "content": "<|user|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151337": {
+      "content": "<|assistant|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151338": {
+      "content": "<|observation|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151339": {
+      "content": "<|begin_of_image|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151340": {
+      "content": "<|end_of_image|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151341": {
+      "content": "<|begin_of_video|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151342": {
+      "content": "<|end_of_video|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151343": {
+      "content": "<|begin_of_audio|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151344": {
+      "content": "<|end_of_audio|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151345": {
+      "content": "<|begin_of_transcription|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151346": {
+      "content": "<|end_of_transcription|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151347": {
+      "content": "<|code_prefix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151348": {
+      "content": "<|code_middle|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151349": {
+      "content": "<|code_suffix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151350": {
+      "content": "<think>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151351": {
+      "content": "</think>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151352": {
+      "content": "<tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151353": {
+      "content": "</tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151354": {
+      "content": "<tool_response>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151355": {
+      "content": "</tool_response>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151356": {
+      "content": "<arg_key>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151357": {
+      "content": "</arg_key>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151358": {
+      "content": "<arg_value>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151359": {
+      "content": "</arg_value>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151360": {
+      "content": "/nothink",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151361": {
+      "content": "<|begin_of_box|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151362": {
+      "content": "<|end_of_box|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151363": {
+      "content": "<|image|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151364": {
+      "content": "<|video|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [
+    "<|endoftext|>",
+    "[MASK]",
+    "[gMASK]",
+    "[sMASK]",
+    "<sop>",
+    "<eop>",
+    "<|system|>",
+    "<|user|>",
+    "<|assistant|>",
+    "<|observation|>",
+    "<|begin_of_image|>",
+    "<|end_of_image|>",
+    "<|begin_of_video|>",
+    "<|end_of_video|>",
+    "<|begin_of_audio|>",
+    "<|end_of_audio|>",
+    "<|image|>",
+    "<|video|>",
+    "<|begin_of_transcription|>",
+    "<|end_of_transcription|>",
+    "<|code_prefix|>",
+    "<|code_middle|>",
+    "<|code_suffix|>",
+    "/nothink"
+  ],
+  "clean_up_tokenization_spaces": false,
+  "do_lower_case": false,
+  "eos_token": "<|endoftext|>",
+  "extra_special_tokens": {},
+  "model_max_length": 128000,
+  "pad_token": "<|endoftext|>",
+  "padding_side": "left",
+  "remove_space": false,
+  "tokenizer_class": "PreTrainedTokenizerFast"
+}

video_preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,11 @@

+{
+    "size": {"shortest_edge": 12544, "longest_edge": 100352000},
+    "do_rescale": true,
+    "patch_size": 14,
+    "temporal_patch_size": 2,
+    "merge_size": 2,
+    "image_mean": [0.48145466, 0.4578275, 0.40821073],
+    "image_std": [0.26862954, 0.26130258, 0.27577711],
+    "video_processor_type": "Glm46VVideoProcessor",
+    "processor_class": "Glm46VProcessor"
+}