JessYanCoding / JessYanCoding/hf-trending-mirror

trending-data (auto-updated daily; do not close)

Open
#2 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

{"fetchedAt": "2026-05-26T22:55:38.084115Z", "models": [{"id": "bytedance-research/Lance", "pipeline_tag": "any-to-any", "downloads": 1908, "likes": 858, "last_modified": "2026-05-26T17:40:11.000Z", "tags": ["Lance", "safetensors", "multimodal", "image-generation", "video-generation", "image-editing", "video-understanding", "any-to-any", "arxiv:2605.18678", "base_model:Qwen/Qwen2.5-VL-3B-Instruct"], "readme": "---\nlicense: apache-2.0\nbase_model:\n- Qwen/Qwen2.5-VL-3B-Instruct\npipeline_tag: any-to-any\nlibrary_name: Lance\ntags:\n- multimodal\n- image-generation\n- video-generation\n- image-editing\n- video-understanding\n- any-to-any\n---\n\n

\n \"Lance\n\n

\n Lance: Unified Multimodal Modeling by Multi-Task Synergy\n

\n\n

\n \n Fengyi Fu*,\n Mengqi Huang*,✉,\n Shaojin Wu*,\n Yunsheng Jiang*,\n Yufei Huo,\n Jianzhu Guo✉,§\n \n
\n\n \n Hao Li,\n Yinghang Song,\n Fei Ding,\n Qian He,\n Zheren Fu,\n Zhendong Mao,\n Yongdong Zhang\n \n
\n ByteDance\n
\n * Equal contribution   \n Corresponding authors   \n § Project lead\n

\n

\n \"Homepage\"\n \n\n

 Marlin: a tiny VLM to extract structured information from videos

\n
\n\nMarlin is a 2B video VLM tuned for the two questions developers actually like ask their videos: **what** is happening, and **when?** It produces structured Scene + Event captions with second-precise timestamps, and resolves natural-language queries to span-grounded (start, end) ranges in the video. At 2B params, it is the strongest open model in its weight class on dense captioning (DREAM-1K, CaReBench) and natural-language temporal grounding (TimeLens-Bench), and competitive with Gemini-2.5 at a fraction of the cost.\n\n## ✨ Key features\n\n- 📝 **State-of-the-art dense captioning at 2B.** Tops the CaReBench leaderboard and sits between Tarsier-34B and Gemini-1.5-Pro on DREAM-1K, two of the most rigorous fine-grained video-captioning benchmarks in the community.\n- ⏱️ **Best-in-class temporal grounding at 2B.** On Tencent's TimeLens-Bench (Charades / ActivityNet / QVHighlights), Marlin beats Qwen2.5-VL-7B by +6.4 mIoU and matches Gemini-2.0-Flash.\n- 🔥 **Built to deploy.** 2B params, vLLM- and swift-deploy-compatible, runs on a single consumer GPU. Same canonical training prompt at inference time, no special wrappers required.\n- 🛠️ **Developer-friendly.*", "params_total": null}, {"id": "meituan-longcat/LongCat-Video-Avatar-1.5", "pipeline_tag": null, "downloads": 0, "likes": 296, "last_modified": "2026-05-26T03:21:29.000Z", "tags": ["diffusers", "onnx", "safetensors", "audio-text-to-video", "audio-image-text-to-video", "audio-driven-video-continuation", "transformers", "avatar", "video-generation", "en"], "readme": "---\nlicense: mit\nlanguage:\n- en\n- zh\nlibrary_name: diffusers\ntags:\n- audio-text-to-video\n- audio-image-text-to-video\n- audio-driven-video-continuation\n- diffusers\n- transformers\n- avatar\n- video-generation\n---\n# LongCat-Video-Avatar-1.5\n\n
\n \"LongCat-Video\"\n
\n
\n\n
\n \n \n \n
\n\n
\n \n \n
\n\n
\n \n
\n\n## 🚀 Model Introduction\nWe are excited to announce the release of LongCat-Video-Avatar 1.5, an upgraded open-source framework that prioritizes extreme empirical optimization and production-readiness for audio-driven human video generation. Built upon the LongCat-Video foundation model, v1.5 delivers highly stable, commercial-grade avatar video synthesis supporting native tasks including Audio-Text-to-Video (AT2V), Audio-Text-Image-to-Video (ATI2V), and Video Continuation, with seamless compatibility for both single-stream and multi-stream audio inputs.\n\n### Key Features\n-", "params_total": null}, {"id": "openbmb/MiniCPM5-1B", "pipeline_tag": "text-generation", "downloads": 2409, "likes": 301, "last_modified": "2026-05-26T04:24:04.000Z", "tags": ["transformers", "safetensors", "llama", "text-generation", "minicpm", "minicpm5", "long-context", "tool-calling", "on-device", "edge-ai"], "readme": "---\nlicense: apache-2.0\nlanguage:\n- en\n- zh\nlibrary_name: transformers\npipeline_tag: text-generation\ntags:\n- minicpm\n- minicpm5\n- llama\n- text-generation\n- long-context\n- tool-calling\n- on-device\n- edge-ai\ndatasets:\n- openbmb/Ultra-FineWeb\n- openbmb/Ultra-FineWeb-L3\n- openbmb/UltraData-Math\n- openbmb/UltraData-SFT-2605\n---\n\n
\n\n
\n\n

\nMiniCPM Tech Report |\nGitHub Repo |\nUltraData |\nMiniCPM Desk Pet |\nOnline Demo\n

\n\n

\nEnglish |\n中文\n

\n\n## Highlights\n\nWe are releasing **MiniCPM5-1B**, the first model in the **MiniCPM5** series. It is a dense 1B Transformer built for on-device, local deployment, and resource-constrained scenarios, reaching 1B-class open-source SOTA.\n\n🏆 **1B-class open-source SOTA**: compared with strong open-source models in the same size class, MiniCPM5-1B reaches SOTA within this comparison set. Its advantage is most visible in agentic tool use, code generation, and difficult reasoning.\n\n![MiniCPM5-1B capability comparison by domain](https://raw.githubusercontent.com/OpenBMB/MiniCPM/main/assets/minicpm5/public_leaderboard_radar_en.png)\n\n🧠 **Hybrid Reasoning**: built-in `` chat template, switch via `enable_thinking`. The same checkpoint serves as both a fast assistant and a deliberate reasoner.\n\n🛠️ **Deployment / Fine-tuning Resources**: the MiniCPM GitHub repo provides single-page cookbooks and Agent Skills for major inference backends and fine-tu", "params_total": null}, {"id": "sapientinc/HRM-Text-1B", "pipeline_tag": "text-generation", "downloads": 103033, "likes": 377, "last_modified": "2026-05-21T05:57:38.000Z", "tags": ["transformers", "safetensors", "hrm_text", "text-generation", "hrm", "hierarchical-reasoning", "prefix-lm", "pre-alignment", "non-chat", "non-instruction-tuned"], "readme": "---\nlicense: apache-2.0\nlanguage:\n- en\nlibrary_name: transformers\npipeline_tag: text-generation\ntags:\n- hrm\n- hierarchical-reasoning\n- prefix-lm\n- pre-alignment\n- non-chat\n- non-instruction-tuned\n---\n\n![HRM-Text banner](banner.jpg)\n\n![Benchmark scatter: FLOPs and tokens vs benchmark average for HRM-Text-1B vs comparable models](benchmark_scatter.png)\n\n

\n \"arXiv\n \"GitHub\"\n

\n\n# HRM-Text-1B\n\nA 1 B-parameter language model checkpoint built on the **Hierarchical Reasoning Model (HRM)** architecture, trained by Sapient Intelligence from scratch on structured public datasets. \n\nHRM is a dual-timescale recurrent architecture: two Transformer modules (H = high-level / slow, L = low-level / fast) iterate over the same input embeddings for `H_cycles × (L_cycles + 1)` steps, with additive state injection (`z_L + z_H`). This gives effectively unbounded compute depth at bounded parameter count.\n\n## Disclaimer\n\nThis is a **pre-alignment** model checkpoint, not a chat or instruction-following assistant. It is pre-trained on a PrefixLM objective with condition prefix tokens and has **not** been multi-turn dialogue tuned, long-context adapted, instruction-tuned, RLHF-trained, or otherwise aligned for assistant-style use. If you want to use HRM-Text like a chat model, you would need to perform further alignment, such as SFT and/or RL, on task-specific data. This checkpoint is meant to serve as a starting point, not a finished assistant.\n\nPractical guidance for prompting the raw checkpoint:\n\n- **NLP tasks (classification, extraction, structured output, short-form QA)**: use the `direct` condition with 2–8 few-shot in-context examples. `direct` + few-shot is the str", "params_total": null}, {"id": "Supertone/supertonic-3", "pipeline_tag": "text-to-speech", "downloads": 48112, "likes": 695, "last_modified": "2026-05-18T08:59:01.000Z", "tags": ["supertonic", "onnx", "text-to-speech", "speech-synthesis", "tts", "multilingual", "on-device", "en", "ko", "ja"], "readme": "---\nlicense: openrail\nlanguage:\n- en\n- ko\n- ja\n- ar\n- bg\n- cs\n- da\n- de\n- el\n- es\n- et\n- fi\n- fr\n- hi\n- hr\n- hu\n- id\n- it\n- lt\n- lv\n- nl\n- pl\n- pt\n- ro\n- ru\n- sk\n- sl\n- sv\n- tr\n- uk\n- vi\npipeline_tag: text-to-speech\ntags:\n- text-to-speech\n- speech-synthesis\n- tts\n- onnx\n- multilingual\n- on-device\nlibrary_name: supertonic\n---\n\n# Supertonic 3 | Lightning Fast, On-Device, Accurate TTS\n\n![Supertonic 3 Preview](img/Supertonic3_HeroImage.png)\n\n

\n \"Demo\"\n \"Code\"\n \"Python\n

\n\n**Supertonic** is a lightweight text-to-speech system for local inference. It runs with ONNX Runtime entirely on your device, with no cloud call required for synthesis.\n\n**Supertonic 3** expands the open-weight release from 5 to **31 languages**, improves reading stability, and reduces repeat/skip failures.\n\n## Quick Start\n\nInstall the Python SDK and generate speech immediately. On first run, the SDK downloads the model assets from Hugging Face.\n\n```bash\npip install supertonic\n```\n\n```python\nfrom supertonic import TTS\n\ntts = TTS(auto_download=True)\nstyle = tts.get_voice_style(voice_name=\"M1\")\n\ntext = \"A gentle breeze moved through the open window while everyone listened to the story.\"\nwav, duration = tts.synthesize(text, voice_style=style, lang=\"en\")\n\ntts.save_audio(wav, \"output.wav\")\nprint(f\"Generated {duration:.2f}s of audio\")\n```\n\n## What's New in Supertonic 3\n\n- **31 languages**: expanded from the 5-language Supertonic 2 release.\n- **More stable reading**: fewer repeat and skip failures, especially on short and long utterances", "params_total": null}, {"id": "HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive", "pipeline_tag": "image-text-to-text", "downloads": 1598473, "likes": 906, "last_modified": "2026-04-17T02:59:21.000Z", "tags": ["gguf", "uncensored", "qwen3.6", "moe", "vision", "multimodal", "image-text-to-text", "en", "zh", "multilingual"], "readme": "---\nlicense: apache-2.0\ntags:\n- uncensored\n- qwen3.6\n- moe\n- gguf\n- vision\n- multimodal\nlanguage:\n- en\n- zh\n- multilingual\npipeline_tag: image-text-to-text\nbase_model: Qwen/Qwen3.6-35B-A3B\n---\n\n# Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive\n\n> **[Join the Discord](https://discord.gg/SZ5vacTXYf)** for updates, roadmaps, projects, or just to chat.\n\nQwen3.6-35B-A3B uncensored by HauhauCS. **0/465 Refusals.**\n\n> **HuggingFace's \"Hardware Compatibility\" widget doesn't recognize K_P quants** — it may show fewer files than actually exist. Click **\"View +X variants\"** or go to **Files and versions** to see all available downloads.\n\n## About\n\nNo changes to datasets or capabilities. Fully functional, 100% of what the original authors intended - just without the refusals.\n\nThese are meant to be the best lossless uncensored models out there.\n\n## Aggressive Variant\n\nStronger uncensoring — model is fully unlocked and won't refuse prompts. May occasionally append short disclaimers (baked into base model training, not refusals) but full content is always generated.\n\nFor a more conservative uncensor that keeps some safety guardrails, check the Balanced variant when it's available.\n\n## Downloads\n\n| File | Quant | BPW | Size |\n|------|-------|-----|------|\n| [Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q8_K_P.gguf](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive/resolve/main/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q8_K_P.gguf) | Q8_K_P | 10.06 | 44 GB |\n| — | Q8_0 | 8.5 | — |\n| [Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive/resolve/main/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf) | Q6_K_P | 7.07 | 31 GB |\n| — | Q6_K | 6.6 | — |\n| [Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q5_K_P.gguf](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive/resolve/main/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q5_K_P.gg", "params_total": null}, {"id": "CohereLabs/command-a-plus-05-2026-w4a4", "pipeline_tag": "image-text-to-text", "downloads": 7769, "likes": 206, "last_modified": "2026-05-22T14:39:44.000Z", "tags": ["transformers", "safetensors", "cohere2_vision", "image-text-to-text", "conversational", "chat", "en", "ar", "bg", "bn"], "readme": "---\ninference: false\nlibrary_name: transformers\nlanguage:\n- en\n- ar\n- bg\n- bn\n- ca\n- cs\n- da\n- de\n- el\n- es\n- et\n- fa\n- fi\n- fil\n- fr\n- ga\n- he\n- hi\n- hr\n- hu\n- id\n- is\n- it\n- ja\n- ko\n- lt\n- lv\n- ms\n- mt\n- nl\n- 'no'\n- pa\n- pl\n- pt\n- ro\n- ru\n- sk\n- sl\n- sr\n- sv\n- ta\n- te\n- th\n- tr\n- uk\n- ur\n- vi\n- zh\nlicense: apache-2.0\nbase_model: CohereLabs/command-a-plus-05-2026\nbase_model_relation: quantized\npipeline_tag: image-text-to-text\ntags:\n- conversational\n- chat\n---\n\n# **Model Card for Command A+**\n\n## **Model Summary**\n\nCommand A+ is an open source model with 25 billion active parameters and 218B total parameters model optimized for agentic, multilingual, and reasoning-heavy tasks with a focus on enterprise performance, while also providing support for vision inputs for processing image inputs.\n\nDeveloped by: [Cohere](https://cohere.com/) and [Cohere Labs](https://cohere.com/research)\n\n* Point of Contact: [**Cohere Labs**](https://cohere.com/research)\n* License: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)\n* Model: command-a-plus-05-2026\n* Model Size: 25B active parameters, 218B total parameters\n* Context length: 128K input\n\nFor more details about this model, please check out our [blog post](http://cohere.com/blog/command-a-plus).\n\nYou can try out Command A+ before downloading the weights in our hosted [Hugging Face Space](https://huggingface.co/spaces/CohereLabs/command-a-plus-05-2026).\n\n**Available quantizations**\n\nThe following quantizations are available with example minimum GPU requirements\n\n| Quantization | Blackwell | Hopper |\n| :---- | :---- | :---- |\n| [BF16 (16-bit)](https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16) | 4 x B200 | 8 x H100 |\n| [FP8 (8-bit)](https://huggingface.co/CohereLabs/command-a-plus-05-2026-fp8) | 2 x B200 | 4 x H100 |\n| [W4A4 (4-bit)](https://huggingface.co/CohereLabs/command-a-plus-05-2026-w4a4) | 1 x B200 | 2 x H100 |\n\nAll three quantizations show negligible differences in benchmark quality and performance. **Ou", "params_total": null}, {"id": "SulphurAI/Sulphur-2-base", "pipeline_tag": "text-to-video", "downloads": 1376847, "likes": 1372, "last_modified": "2026-05-22T01:12:13.000Z", "tags": ["diffusers", "gguf", "text-to-video", "endpoints_compatible", "region:us", "conversational"], "readme": "---\nlibrary_name: diffusers\npipeline_tag: text-to-video\n---\n\n**Sulphur 2**\n\nAn uncensored video generation model based on LTX 2.3 supporting both t2v and i2v natively, as well as all of the other ltx 2.3 formats.\n\nJoin our **[Discord](https://discord.gg/GSXJhKZ9V)**\n\nSupport the next version of the project, even just a few dollars would go a long way: **[Kofi](https://ko-fi.com/fusioncow)**\n\n---\n\n**Get Started:**\nTo get started with the model, I recommend downloading either of the dev versions, (fp8mixed or bf16) and downloading the distill lora provided. By the way, I'm aware the workflows contain sulphur_final right now, just use the lora or use the full models, don't use both at the same time.\n\nThis model contains a **prompt enhancer**. The easiest way to get started with the prompt enhancer is by using it on lmstudio. The way to accomplish this is by going to your model folder inside lmstudio, then opening it up in your file explorer. Create a folder named \"Sulphur\", then a folder inside that called \"promptenhancer\". Inside that folder, place the gguf file and the mmproj file. Once you've done that, you should be able to load the prompt enhancer in lmstudio. There is no system prompt for it, just send the text (and an image) you'd like to be enhanced.\n\n*As a note, this readme will contain better setup instructions and how to train on top of the model soon.\n\n---\n\n**Links**\n- **([CivitAI Base Model](https://civitai.red/models/2594061/sulphur-2-base))** -\n- **([CivitAI Quant Model](https://civitai.red/models/2630742))** -\n\n\n**Credits**\n\n- **([TenStrip](https://huggingface.co/TenStrip))** — Testing & model merging ([His i2v merge of sulphur 2, highly recommend for i2v](https://huggingface.co/TenStrip/LTX2.3-10Eros))\n- **@s1lv3rc01n** — Testing & model merging/quantizing ([silveroxides](https://huggingface.co/silveroxides))\n- **@mov7162** — Musubi Tuner guidance\n- And many others, if you'd like to be on the credits and I didn't place you here, message me I likely as", "params_total": null}, {"id": "deepseek-ai/DeepSeek-V4-Pro", "pipeline_tag": "text-generation", "downloads": 5019884, "likes": 4308, "last_modified": "2026-05-06T04:18:44.000Z", "tags": ["transformers", "safetensors", "deepseek_v4", "text-generation", "conversational", "license:mit", "eval-results", "endpoints_compatible", "8-bit", "fp8"], "readme": "---\nlicense: mit\nlibrary_name: transformers\n---\n# DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence\n\n\n\n\n\n
\n \"DeepSeek-V4\"\n
\n
\n
\n \n \"Homepage\"\n \n \n \"Chat\"\n \n
\n
\n \n \"Hugging\n \n \n \"Twitter\n \n
\n
\n \n \"License\"\n \n
\n\n

\n Technical Report👁️\n

\n\n## Introduction\n\nWe present a pre", "params_total": null}, {"id": "unsloth/Qwen3.6-27B-MTP-GGUF", "pipeline_tag": "image-text-to-text", "downloads": 735349, "likes": 502, "last_modified": "2026-05-26T08:45:52.000Z", "tags": ["transformers", "gguf", "unsloth", "qwen", "qwen3_5", "image-text-to-text", "base_model:Qwen/Qwen3.6-27B", "base_model:quantized:Qwen/Qwen3.6-27B", "license:apache-2.0", "endpoints_compatible"], "readme": "---\nlibrary_name: transformers\nlicense: apache-2.0\nlicense_link: https://huggingface.co/Qwen/Qwen3.6-27B/blob/main/LICENSE\npipeline_tag: image-text-to-text\nbase_model:\n- Qwen/Qwen3.6-27B\ntags:\n- unsloth\n- qwen\n- qwen3_5\n---\n# Read our How to [Run Qwen3.6 MTP Guide!](https://unsloth.ai/docs/models/qwen3.6#mtp-guide)\n
\n

\n See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks.\n

\n
\n \n \n \n \n \n \n \n \n \n
\n
    \n
  • MTP enables ~1.5-2x faster inference with no accuracy loss.\n
  • You can now run Qwen3.6 MTP GGUFs in Unsloth Studio\n
  • Unsloth Studio auto sets the ideal MTP settings for your hardware (Mac, CPU, GPU):\n\n\"qwen3.6\n\n### To run in llama.cpp:\n\n```bash\napt-get update\napt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y\ngit clone https://github.com/ggml-org/llama.cpp\ncmake llama.cpp -B llama.cpp/build \\\n -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON\ncmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-mtmd-cli llama-server llama-gguf-split\ncp llama", "params_total": null}, {"id": "openbmb/MiniCPM-V-4.6", "pipeline_tag": "image-text-to-text", "downloads": 314347, "likes": 978, "last_modified": "2026-05-19T14:01:40.000Z", "tags": ["transformers", "safetensors", "minicpmv4_6", "image-text-to-text", "minicpm-v", "multimodal", "On-Device Model", "lightweight", "conversational", "arxiv:2604.27393"], "readme": "---\nlicense: apache-2.0\npipeline_tag: image-text-to-text\ntags:\n- minicpm-v\n- multimodal\n- On-Device Model\n- lightweight\nlibrary_name: transformers\n---\n\nA Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone\n\n[GitHub](https://github.com/OpenBMB/MiniCPM-o) | [CookBook](https://github.com/OpenSQZ/MiniCPM-V-CookBook) | [Demo](https://huggingface.co/spaces/openbmb/MiniCPM-V-4.6-Demo) |\n[Feishu (Lark)](https://raw.githubusercontent.com/openbmb/MiniCPM-V/main/assets/feishu_qrcode.png)\n\n## News\n\n* [2026.05.17] ⭐️⭐️⭐️ We release the API service of MiniCPM-V 4.6, with a **public free API key** together! Try [it](https://github.com/OpenBMB/MiniCPM-V/blob/main/docs/api.md) now.\n\n\n\n## MiniCPM-V 4.6\n\n**MiniCPM-V 4.6** is our most edge-deployment-friendly model to date. The model is built based on SigLIP2-400M and the Qwen3.5-0.8B LLM. It inherits the strong single-image, multi-image, and video understanding capabilities of MiniCPM-V family, while significantly improving computation efficiency. It also introduces mixed 4x/16x visual token compression. Notable features of MiniCPM-V 4.6 include:\n\n- 🔥 **Leading Foundation Capability.**\n MiniCPM-V 4.6 scores 13 on the Artificial Analysis Intelligence Index benchmark, outperforming Qwen3.5-0.8B's score of 10 with 19x fewer token cost, and Qwen3.5-0.8B-Thinking's score of 11 with 43x fewer token cost. It also surpasses the larger Ministral 3 3B (score of 11).\n\n- 💪 **Strong Multimodal Capability.**\n MiniCPM-V 4.6 outperforms Qwen3.5-0.8B on most vision-language understanding tasks, and reaches Qwen3.5 2B-level capability on many benchmarks including OpenCompass, RefCOCO, HallusionBench, MUIRBench, and OCRBench.\n- 🚀 **Ultra-Efficient Architecture.**\n Based on the latest technique in [LLaVA-UHD v4](https://github.com/THUMAI-Lab/LLaVA-UHD-v4), MiniCPM-V 4.6 reduces the visual encoding computation FLOPs by more than 50%. It enables MiniCPM-V 4.6 to achieve better efficiency to even smaller models, achieving ~1", "params_total": null}, {"id": "Jackrong/Qwopus3.6-27B-v2-GGUF", "pipeline_tag": "image-text-to-text", "downloads": 16379, "likes": 143, "last_modified": "2026-05-24T08:15:12.000Z", "tags": ["transformers", "gguf", "text-generation-inference", "image", "unsloth", "qwen3_6", "reasoning", "chain-of-thought", "lora", "sft"], "readme": "---\nbase_model:\n- qwen/Qwen3.6-27B\ntags:\n- text-generation-inference\n- image\n- transformers\n- unsloth\n- qwen3_6\n- reasoning\n- chain-of-thought\n- lora\n- sft\n- multimodal\n- vision\n- tool-use\n- function-calling\n- long-context\n- agent\n- science\nlicense: apache-2.0\nlanguage:\n- en\n- zh\n- es\n- ru\n- ja\npipeline_tag: image-text-to-text\ndatasets:\n- Jackrong/Claude-opus-4.6-TraceInversion-9000x\n- Jackrong/Claude-opus-4.7-TraceInversion-5000x\n---\n\n
    \n
    \n
    \n

    🪐 Qwopus3.6-27B-v2

    \n SFT Release\n
    \n

    Reasoning-Enhanced Dense Language Model Fine-Tuned on Qwen3.6-27B

    \n
    \n
    \n 🧬 Trace Inversion & Negentropy\n 🧠 27B Parameters\n ", "params_total": null}, {"id": "numind/NuExtract3", "pipeline_tag": "image-to-text", "downloads": 20350, "likes": 160, "last_modified": "2026-05-20T09:24:47.000Z", "tags": ["transformers", "safetensors", "image-text-to-text", "qwen3_5", "vision-language", "vlm", "document-understanding", "structured-extraction", "information-extraction", "ocr"], "readme": "---\nlicense: apache-2.0\nlicense_link: https://huggingface.co/numind/NuExtract3/blob/main/LICENSE\nlibrary_name: transformers\npipeline_tag: image-to-text\ntags:\n- image-text-to-text\n- transformers\n- safetensors\n- qwen3_5\n- vision-language\n- vlm\n- document-understanding\n- structured-extraction\n- information-extraction\n- ocr\n- document-to-markdown\n- markdown\n- rag\n- reasoning\n- multilingual\n- conversational\nbase_model:\n- Qwen/Qwen3.5-4B\nmodel_name: NuExtract3\n---\n\n

    \n \n \n \n

    \n\n\n

    \n 🖥️ API / Platform   |   \n 📑 Blog   |   \n 🗣️ Discord   |   \n 🛠️ GitHub\n

    \n\n**NuExtract3** is a unified **4B** vision-language reasoning model for document understanding.\n\nIt combines strong **structured information extraction** with high-quality **image-to-Markdown** conversion, making it suitable for extraction pipelines, OCR, and RAG preprocessing for all types of documents such as scans, receipts, forms, invoices, contracts or tables.\n\nTry it out in [the 🤗 space!](https://huggingface.co/spaces/numind/NuExtract-3-4B)\n\n## Overview\n\n- **Structured extraction**: input (text/images) + JSON template + instructions --> JSON output\n- **Markdown conversion**: input (text/images) --> Markdown\n- **Multimodal inputs**: text, images, or text + images.\n- **Multilingual** documents.\n- **Reasoning** and non-reasoning inference modes.\n- **Template generation** for structured extraction from natural language or input document.\n\n# Benchmark results\n\n## Structured Extraction\n\nWe benchmarked NuExtract on NuMind's internal structured benchmark, measuring model's performances on ~600 documents of diverse types including invoices, movie poster", "params_total": null}, {"id": "circlestone-labs/Anima", "pipeline_tag": null, "downloads": 676447, "likes": 1556, "last_modified": "2026-05-14T17:13:39.000Z", "tags": ["diffusion-single-file", "comfyui", "license:other", "region:us"], "readme": "---\nlicense: other\nlicense_name: circlestone-labs-non-commercial-license\nlicense_link: LICENSE.md\ntags:\n- diffusion-single-file\n- comfyui\n---\n\n\n\nAnima is a 2 billion parameter text-to-image model created via a collaboration between CircleStone Labs and Comfy Org. It is focused mainly on anime concepts, characters, and styles, but is also capable of generating a wide variety of other non-photorealistic content. The model is designed for making illustrations and artistic images, and will not work well at realism.\n\nIt is trained on several million anime images and about 800k non-anime artistic images. No synthetic data was used for training. The knowledge cut-off date for the anime training data is September 2025.\n\n**NEW:** Try the [Turbo LoRA](https://civitai.com/models/2560840/anima-turbo-lora) for better stability and much faster generations.\n\n# Versions\n- Anima-Base\n - The pretrained, unrefined base model. Maximum flexibility, diversity, and style adherence.\n- Anima-Turbo\n - Coming soon.\n\n# Installing and running\nWorkflow:\n\nThe model is natively supported in ComfyUI. The above image contains a workflow; you can open it in ComfyUI or drag-and-drop to get the workflow. The model files go in their respective folders inside your model directory:\n- anima-base-v1.0.safetensors goes in ComfyUI/models/diffusion_models\n- qwen_3_06b_base.safetensors goes in ComfyUI/models/text_encoders\n- qwen_image_vae.safetensors goes in ComfyUI/models/vae (this is the Qwen-Image VAE, you might already have it)\n\n## Generation settings\n- Works at resolutions between 512^2 and 1536^2 pixels.\n- 30-50 steps, CFG 4-5.\n- A variety of samplers work. Some of my favorites:\n - er_sde: neutral style, flat colors, sharp lines. I use this as a reasonable default.\n - euler_a: Softer, thinner lines. Can sometimes tend towards a 2.5D look. CFG can be pushed a bit higher than other samplers without burning the image.\n - dpmpp_2m_sde_gpu: similar", "params_total": null}]}

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.