Accuracy-Efficiency
Text / LLMIt predicts labels for new rows in tabular data for classification and regression tasks. The model is available in three sizes from 28M to 215M parameters.
📡 Daily radar
What came out that you can actually download — grouped by what it makes, with the specs that decide whether it runs on your box. Every link goes to the weights, the repository or the paper, never to coverage of them. Pick a date to read an earlier day, or a modality to see only what it shipped.
Archive
Modality
6 releases on Wednesday, 30 September 2026
6 releases226 clusters from 235 items
It predicts labels for new rows in tabular data for classification and regression tasks. The model is available in three sizes from 28M to 215M parameters.
It performs mathematical proof verification and image-to-text processing. It runs using the transformers library and safetensors.
Produces speech and audio from text inputs. Requires 32GB VRAM, an AMD Radeon AI PRO R9700, and ROCm 7.2.
The model performs novel view synthesis via image-to-image generation. The model requires the pytorch library and uses safetensors weights.
The benchmark monitors whether a model's performance declines after it is released. It requires a Claude Max subscription and headless Claude Code.
The model produces ImageNet-1K classification, COCO object detection, and ADE20K semantic segmentation. It requires the Hugging Face CLI to download weights and the code repository for model definitions.
6 releases160 clusters from 169 items
The model produces text output. No hardware requirements are listed because open weights are not available.
Produces text-to-speech audio and voice. Requires 32GB VRAM, an AMD Radeon AI PRO R9700, and ROCm 7.2.
It produces zero-shot voice-cloning and text-to-speech audio for Badini Kurdish. It requires a short reference audio clip and an exact transcript for each generation.
The model performs text generation, vision-language processing, audio, and video understanding. The model is available in 8-bit and fp8 formats.
The model generates text from multimodal inputs including audio, vision, and video. The model is available in 8-bit.
The model performs audio-to-audio voice conversion for Chinese content. Running requires the model weight files and source code.
4 releases84 clusters from 88 items
The model produces zero-shot voice-cloned audio for Badini Kurdish. It requires a reference clip and a transcript for each generation.
It produces diagrams from a text language where the user specifies the placement of elements. The software is released under an Apache-2.0 license.
3 releases97 clusters from 97 items
The Bonsai LLM provides vision and language processing. The 2-billion-parameter model runs on the Snapdragon AR1 Gen 1 Platform.
Produces audio from Japanese text input. The model uses XTTS architecture and is licensed under agpl-3.0.
The repository provides training-content summaries and documentation for model versions. The model features 16B total parameters and 3B active parameters.
8 releases188 clusters from 197 items
Produces vision-language output by generating a block of candidate tokens from shared image and text representations. The model has 3B parameters and supports GGUF, llama.cpp, MLX-VLM, and SGLang.
The 2-billion-parameter model processes vision and language tasks for smart glasses. The model runs on Snapdragon AR1 Gen 1 Platform hardware.
It generates text based on a Qwen3.5-4B fine-tune. The model has 4B parameters and an Apache-2.0 license.
The model performs joint image understanding and generation. The model contains 16B total parameters and 3B active parameters.
Produces Japanese text-to-speech audio. Requires an XTTS inference environment.
The repository provides training-content summaries and documentation. The repository contains documentation, not model weights.
It produces decision model outputs in a single forward pass without token-by-token generation. It runs on a local server via an HTTP API.
The application provides a shared canvas and SDK for humans and agents to architect software. The application runs on macOS and Fedora and integrates with Claude Code, Codex, GPT-6 Sol, and Claude Opus 5.5.
10 releases200 clusters from 210 items
The tokenizer encodes text and estimates token counts for English, Korean, and Japanese. Requires the Transformers library.
The model generates images and performs tasks via custom agent skills. The system runs at 60 FPS on Ubuntu 24.04+ for ARM64.
These models produce text and code. These models run on local hardware.
The model processes video, audio, image, and text for any-to-any tasks. The model contains 31B total parameters, 3B active parameters, and a 256k token context window.
Produces multilingual text-to-speech audio with voice cloning and emotion control. Requires 7.72 GiB of storage and supports 22.05 kHz audio.
It produces text-to-video, video-to-video, image-text-to-video, and video-inpainting using Canny, Depth, HED, MLSD, and Pose control conditions. It runs using the videox_fun library.
Produces image-to-video content using transformer weights, camera adapter, and LoRA modules. Requires the WorldCrafter_Fast directory for shared components including the text encoder, VAE, and scheduler.
Generates image-to-video and text-to-video content with camera control. Requires the WorldCrafter inference code and a pinned uv environment.
The model produces action chunks and denoised video frames from camera frames, robot states, and text instructions. The model contains 7B parameters.
The framework provides tools and building blocks for running local AI. The system requires Apple Silicon hardware and uses the Apache-2.0 license.
12 releases190 clusters from 198 items
The Transformers library enables local inference of quantized models. It uses the GGUF format and supports 27B parameter models.
This tokenizer encodes text and estimates token counts for English, Korean, and Japanese. Running this requires the Transformers library.
Produces image edits and rewritten prompts. 7B parameters.
The model performs any-to-any multimodal tasks with long-context and agentic capabilities. No technical specifications are provided.
The model produces text and scored 28.8% on Terminal-Bench 4.0. The provided text contains no specific hardware requirements.
The model performs any-to-any multimodal tasks for video, audio, image, and text. The model uses 31B total parameters and 3B active parameters for a 256k context window.
Produces multilingual text-to-speech and zero-shot voice cloning. Runs as a 0.6B parameter model with INT4 quantization.
Produces image-to-video and text-to-video content. Requires the diffusers library and a pinned uv environment.
MLX is a framework for running local AI. The software is optimized for Apple Silicon hardware.
The model generates text from multimodal inputs including vision, audio, and video. The model runs using 8-bit or fp8 precision.
Generates images from text prompts. Uses 7B parameters.
The model performs any-to-any tasks including conversational text generation and image-to-text processing. The model runs with the transformers library under the Apache-2.0 license.
14 releases159 clusters from 168 items
Encodes text and estimates token counts for English, Korean, and Japanese. Requires the Transformers library and contains a vocabulary of 196,608 tokens.
This model produces image-to-image transformations and rewritten prompts. The model has 7B parameters.
The model produces text. The source text provides no hardware requirements or model size specifications.
The model performs any-to-any tasks including text generation in multiple languages. The model has 29B parameters and utilizes the transformers library.
Produces yes/no, multiple-choice, and rating outputs for decision-making tasks. Available in 0.8B, 4B, and 9B sizes for CUDA, ROCm, and Apple Silicon hardware.
The model generates text, video, audio, and image content. The model uses 31 billion total parameters and 3 billion active parameters per token.
The model produces text-to-speech audio as an observer head for the CosyVoice2-0.5B native-token interface. The model requires the PyTorch library and has 2,892,653 parameters.
This model converts text into multilingual audio and performs zero-shot voice cloning. The model features 0.6B parameters and supports 44.1 kHz audio in INT4 format.
Produces text for multimodal, vision-language, audio, and video-understanding tasks. Runs using 8-bit or fp8 precision with the transformers library.
The model performs text generation and processes multimodal inputs including vision, audio, and video. The model supports 8-bit and fp8 quantization and runs with the transformers library.
This model generates images from text prompts. The model has 7B parameters.
This model generates images from text input. It runs via the transformers library under an Apache-2.0 license.
AIVORENCE published Mia-v1-E2B to the Hugging Face hub for any-to-any. Specs: Apache-2.0.
convaiinnovations published laya-multilingual to the Hugging Face hub for text-classification. Specs: Apache-2.0.
12 releases84 clusters from 92 items
This model performs image editing and prompt rewriting. The model contains 7B parameters.
Produces text for any-to-any tasks in 140 languages with a 128K context window. Requires 26B total parameters with 4B active parameters.
Pirate Face mirrors Hugging Face models as peer-to-peer torrents with SHA-256 verification. Apache-2.0.
This model produces text-to-speech audio using an observer head for the CosyVoice2-0.5B system. The system requires PyTorch and uses 2,892,653 parameters.
Produces text across 166 languages for any-to-any tasks. Requires 7B parameters and INT4 quantization.
The model produces zero-shot text-to-speech and voice cloning in Chinese, English, Japanese, Spanish, and Arabic with emotion control. The model has 0.8B parameters and requires 6GB VRAM to run at a 22.05 kHz sample rate.
Produces text-to-speech audio and clones voices into 17 languages from 6-second clips. The model contains 7B parameters and supports 24 kHz audio.
It produces multi-shot scenes as one continuous take by chaining 10-15 second video blocks with consistent audio and color. The system requires the minimax-h3 library and a ComfyUI node pack.
The model generates 2048x2048 images with alpha channels, rendered text, and processes up to 10 input images for editing. The model runs on 7B parameters using an MMDiT architecture on consumer cards.
This model generates images from text prompts. The model has 7B parameters.
Produces any-to-any outputs including conversational text and image-to-text. Requires the transformers library and safetensors format under an Apache-2.0 license.
The model performs text classification for routing, scoring, and moderation. It runs using the transformers library under the Apache-2.0 license.
8 releases100 clusters from 114 items
The model produces probability predictions over structured schemas using a non-autoregressive architecture and reinforcement learning. The model has 8B parameters, an Apache-2.0 license, and supports 100 languages.
The model performs any-to-any text generation in 166 languages. The model has 7B parameters and uses INT4 quantization.
It produces zero-shot text-to-speech with voice cloning and emotion control in Chinese, English, Japanese, Spanish, and Arabic. The model requires 0.8B parameters and 6GB VRAM to produce 22.05 kHz audio.
Produces text-to-speech and voice cloning in 17 languages from 6-second audio samples. Uses 7B parameters to output 24 kHz audio.
This model converts graphemes to phonemes for text-to-speech. It is a .tflite file with fixed-length [1, 96] and FP32 precision for CPU.
The model generates video from image inputs. No hardware specifications are provided.
The update adds thinking controls to the show API and adds support for Nemotron H vision models on Apple Silicon. Nemotron H vision models run on Apple Silicon via MLX.
The model performs text classification. The model requires the transformers library and uses the safetensors format.
6 releases197 clusters from 210 items
It functions as a large language model. The model has 4B parameters and is licensed under Apache-2.0.
This multimodal model performs reasoning, coding, vision, and agentic tasks. The 27B parameter model requires 5.9GB of memory and supports a 262K context window.
This model converts graphemes to phonemes for the Kokoro text-to-speech system. The model runs as a TFLite file in FP32 format on CPU.
Produces cloned voices in 17 languages from 6-second audio clips. The model has 7 billion parameters and operates at a 24 kHz sample rate.
This library manages memory allocation, provides bug fixes, and reports pinned-memory usage statistics. No hardware or software requirements are listed in the source text.
It indexes visited web pages and local files for search via a web interface, terminal, or AI assistant. Requires a downloaded binary, a running terminal server, and a browser extension for Chrome or Firefox.
6 releases185 clusters from 196 items
The model produces text and handles conversational tasks. The model contains 29B total parameters and 4B active parameters.
Converts graphemes to phonemes for text-to-speech systems. Runs on CPU using the LiteRT framework in FP32 format.
Produces speech in 17 languages and clones voices from 6-second audio samples. Contains 7B parameters and outputs 24 kHz audio.
It generates video sequences based on keyboard inputs and text instructions. The model has 5B parameters and produces output at 24 fps.
Provides 1,271 sentence-aligned pairs translated from Wolof into Modern Standard Arabic for machine translation. Available under a CC-BY-NC license.
It indexes visited web pages and local files for retrieval via a web interface, terminal, or AI assistant connected through MCP. It requires a local binary, a running listener process, and a browser extension for Firefox or Chrome.
4 releases184 clusters from 191 items
The model generates text in English and Chinese. The model has 744B parameters and uses FP8 precision.
Produces speech in 17 languages and clones voices from 6-second audio clips. The model has 7B parameters and 24 kHz audio output.
Produces Turkish speech from text using a speaker-agnostic acoustic model. Requires the antalia library and supports 24 kHz audio.
The software produces images and text through a node-based interface. The system supports a 4TB AMD Windows VA quota.
0 releases169 clusters from 177 items
2 releases135 clusters from 144 items
Produces text in Chinese and English. Requires hardware to run 744B parameters in FP8.
Produces emotionally expressive and duration-controlled zero-shot text-to-speech in English and Chinese. The model uses the safetensors format.
12 releases116 clusters from 122 items
This release provides inference for Hy4-preview, Tencent 770B/49B-active MoE, Qwen3.8-Flash-Next, GraniteSWA, GraniteMoeSWA, and NemotronH_Omni_Reasoning_V3. It supports 770B and 49B parameter models using FP8, BF16, and NVFP4.
Produces text output with a 1M token context window. The model has 780B total parameters and activates 49B parameters per token.
Fine-tunes multiple model families using a declarative system, ternary QAT, and PyTorch 2.13. Qwen3.8-Flash-Next 176.94B MoE requires 120 GiB on a single B300.
Produces text-to-speech audio, speech-editing, and audio-to-audio content. Requires a 1.53B parameter DiT, 24 kHz sampling rate, and 6.76 GB of storage.
The model produces outputs for any-to-any tasks including OCR, speech-recognition, text-to-speech, and video. The model has 5.84 billion parameters and a 2 million token context window.
Produces text-to-speech audio and voice samples. Requires 0.1B parameters with 44.1 kHz sampling and INT8 quantization.
It produces multilingual text-to-speech and zero-shot voice cloning. The model has 0.6B parameters and supports 44.1 kHz audio in INT4 format.
Generates music and voice audio from text input. 3B parameters with a CC-BY-NC-4.0 license.
The model performs feature extraction for audio. The model contains 3B parameters and operates at 48 kHz.
The model produces music representations from 30s audio clips. The model has 632M parameters and 24 layers for 24 kHz mono audio.
The app provides a messaging interface with Material 3 components and large-screen support. No hardware requirements are provided.
It produces text output from image and text inputs for conversational and reasoning tasks. It runs using the transformers library and supports safetensors format under an Apache-2.0 license.
10 releases200 clusters from 206 items
vLLM provides inference for text models including MoE, Qwen, and Nemotron architectures. The software supports FP8, BF16, and NVFP4 precision on CUDA and ROCm hardware.
The model generates text output. The model has 780B parameters, 49B active parameters per token, and a 1M token context window.
It produces fine-tuned language models and multimodal models. It requires a single B300 with 120 GiB of memory to fine-tune a 176.94B parameter model.
This model produces text-to-speech audio. The 1.53B parameter model requires the maestro library.
The model produces text, speech-recognition, and text-to-speech outputs for any-to-any tasks. The model has 5.84 billion parameters and a 2 million token context window.
Produces multilingual text-to-speech audio and voice from text input. The model has 0.1B parameters and supports 44.1 kHz audio in INT8 format.
The model produces multilingual text-to-speech audio and zero-shot voice cloning. The model has 0.6B parameters and supports INT4 quantization.
The model generates music and audio from text inputs. The model has 3 billion parameters.
This model performs feature extraction and functions as an audio autoencoder. The model contains 3B parameters and supports 48 kHz audio.
The model produces feature representations from 30-second audio clips sampled at 24 kHz. The model contains 632 million parameters across 24 layers and 1,024 dimensions.
10 releases208 clusters from 217 items
The model produces text with a 1M token context window. The model requires infrastructure for 780B total parameters and 49B active parameters per token.
The model produces text from image and text inputs. The model has 552B parameters and supports 8-bit precision.
This update implements Model Runner V2 and adds support for models including Hy4-preview, Qwen3.8-Flash-Next, and NemotronH_Omni_Reasoning_V3. It supports 770B and 49B parameter models using FP8, BF16, and NVFP4 formats.
Mistral today announced that it has raised €3 billion in a Series D funding round at a post-money valuation of more than €21 billion. Specs: 3B params.
The model produces text output from image and text inputs. The model has 124B parameters and uses FP8 precision.
The model performs any-to-any tasks including OCR, speech-recognition, text-to-speech, and video. It features 5.84 billion parameters and a 2 million token context window.
It produces music and audio from text. The model has 3 billion parameters.
This model performs feature extraction for audio data using an autoencoder. The model has 3B parameters and supports 48 kHz audio.
Extracts 1,024-dimension music representations from 30-second audio clips. Requires 632 million parameters and 24 kHz mono audio input.
It produces web interfaces. It is a CSS framework.
10 releases113 clusters from 115 items
The model produces text outputs with a 1M token context window. The model contains 780B parameters and activates 49B parameters per token.
The model produces zero-shot time-series forecasts and performs missing value imputation. The model is a 385M-parameter model released under an Apache 2.0 license.
It serves text generation for models using Model Runner V2 and batch-sharded sampling. Supports 770B parameters in FP8 and 49B active MoE configurations.
This large language model produces text. The model has 3 billion parameters.
The model produces any-to-any outputs including OCR, speech recognition, text-to-speech, and video. The model has 5.84 billion parameters and a 2 million token context window.
The model produces text-to-speech audio in English and Chinese. It runs using PyTorch, CUDA, and the transformers library.
This model produces text-to-speech audio for Rioplatense Spanish using zero-shot flow matching and a Vocos vocoder. The model has 19.8M parameters and was trained on 8 hours of Argentinian Spanish.
The model converts text into Eastern Armenian speech. It utilizes a 512-dimensional speaker embedding and 16 kHz audio.
Ollama models can be used in ChatGPT Desktop with support for tool search and response compaction. Setup is available on MacOS and the release includes improvements for Apple Silicon.
Tailwind produces CSS styles for web interface development. The text contains no information regarding hardware or system requirements.
7 releases59 clusters from 61 items
Mistral produces text outputs. The model has 3 billion parameters.
Produces text generation, math, and long-chain-of-thought reasoning. Requires the transformers and pytorch libraries.
The model generates text and executes actions without standard safety safeguards. The model is an open-weight model that is fast and cheap to run.
The model converts Vietnamese text into spoken audio. The model produces 22.05 kHz audio using the Matcha-TTS and Vocos frameworks.
Produces voice and audio from text input. Requires ExecuTorch and a react-native-executorch environment.
Produces Finnish text-to-speech audio. Requires NumPy and Raspberry Pi hardware at 22.05 kHz.
Ollama models run in ChatGPT Desktop with structured output on Apple Silicon. Requires the Ollama app on MacOS.
9 releases156 clusters from 163 items
The model generates text and supports long-context and tool-calling. The model has 2B parameters and an Apache-2.0 license.
The system estimates calorie counts from meal photos and descriptions. The system requires an LLM and the Nutrition5k dataset.
The model produces any-to-any and image-text-to-text content in 140 languages. The model uses 12B total parameters, 4B active parameters, and a 128K context window.
The model produces Finnish text-to-speech audio. It runs on NumPy and supports a 22.05 kHz sample rate.
The model generates speech and audio from text in English and Chinese. The model runs using the transformers library, pytorch, and CUDA.
It converts English text into speech using a FastSpeech2 and HiFi-GAN architecture. The model requires PyTorch and has 59.5 million parameters for 22.05 kHz audio.
The model generates Burmese speech from text input. It requires the f5-tts library and produces 24 kHz audio.
Produces Tamil text-to-speech audio for a single female voice. The 86 MB model uses 4-bit quantised weights to run on a CPU without a GPU.
It integrates Ollama models into ChatGPT Desktop and improves structured output performance on Apple Silicon. Requires the Ollama app on MacOS.
3 releases64 clusters from 71 items
This model produces text and supports tool calling, multi-turn interaction, and error correction. This model requires the transformers library and utilizes the LLaDA2.0-mini architecture.
Produces synchronized audio and video in a single generation pass. No specific hardware or software requirements are listed.
Integrates Ollama models with ChatGPT Desktop and improves structured output on Apple Silicon. Requires the Ollama app on MacOS.
7 releases82 clusters from 93 items
It produces dense and late-interaction embeddings from text tokens and raw image patches for multimodal and multilingual retrieval. Available in 260M and 800M sizes, it runs on an NVIDIA L40S GPU with 2048x2048 image inputs.
The model produces any-to-any and image-text-to-text results. The model uses 26B total parameters with 4B active and a 128K context window.
This model produces text-to-speech audio in English, Japanese, Chinese, and Russian. No hardware or software requirements are listed.
The system produces English speech from text using FastSpeech2 and HiFi-GAN at 22.05 kHz. The model requires PyTorch and contains 59.5 million parameters.
This model converts text into Burmese speech. It requires the f5-tts library and outputs 24 kHz audio.
The model produces multilingual text-to-speech with streaming output and reference-based voice cloning. The model requires the iceAudio library and an active acoustic adapter.
MiniMax H3 produces synchronized audio and video in a single generation pass. The system runs in ComfyUI.
9 releases192 clusters from 204 items
We introduce NeoMME, a family of 260M and 800M multilingual multimodal encoders. Specs: Apache-2.0 · 4K.
OsGo published gemma-4-E2B to the Hugging Face hub for any-to-any. Specs: 26B total / 4B active · 12B params · 128K context · Apache-2.0 · 140 languages.
I'm one of the developers. Specs: 300B params · Apache-2.0 · GGUF.
TimesFM-3 is the third generation of Google Research's zero-shot forecasting model, and the main change from 2.5 is that it handles multivariate inputs natively instead of being limited to a single series' own history. Specs: non-commercial licence.
Available Models Model Compression This section provides transparency about the compression state of each model available on our platform. Specs: 27B params · 8-BIT.
Thanushttk published pocket-tts-hindi to the Hugging Face hub for text-to-speech (voice and audio). Specs: 5s clips · 24 kHz.
jjiang4 published thorsten_vits to the Hugging Face hub for text-to-speech (voice and audio). No specs stated in the source — open the link before quoting numbers.
Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. Specs: GGUF.
Two weeks, one rule, and hundreds of entries from nearly 50 countries. No specs stated in the source — open the link before quoting numbers.
11 releases174 clusters from 194 items
The model produces dense and late-interaction embeddings from text and image patches for visual document retrieval. The model is available in 260M and 800M sizes under the Apache-2.0 license.
The model produces multivariate time series forecasts. The model has 330M parameters and a non-commercial license.
Produces text output using the Qwen 3.8 27B model. The model has 27B parameters and runs at 1500 tokens per second with 8-bit weights.
The model fleet produces text for reasoning, mathematics, coding, and agentic tasks. The suite includes six models with parameter counts ranging from 0.9B to 375B.
The model performs text generation and conversational tasks. The model has 36 billion total parameters and 4 billion active parameters.
The model produces any-to-any outputs in 140 languages. It has 12B parameters and a 128K context window.
The model converts text into Hausa speech audio. It requires the coqui-tts library and XTTS-v2 architecture.
The model performs automatic speech recognition for audio input in 10 languages. The system uses 7B parameters and runs with the transformers library.
The model generates audio output for Qwen and GLM models. The model runs in GGUF format.
Provides a library of WebGPU kernels and an in-browser benchmarking suite. Licensed under Apache-2.0.
Audacity 4.0 provides a clip-editing model, multi-clip selection, grouping, and a split tool. No hardware requirements are provided.
12 releases172 clusters from 185 items
It produces a question-answer dataset based on NCERT textbooks for Indian education. The model requires 8B parameters.
The tool evaluates large language models on benchmark suites. The system runs as a Python-based evaluation harness with ONNX support.
The model produces text from text inputs without vision capabilities. The model requires 770B total parameters, 49B active parameters, a 1M context window, and 1.56TB storage.
Produces text output from text and image inputs across 140 languages. Requires the transformers library to run with 12B parameters and a 128K context window.
The model produces zero-shot text-to-speech audio in multiple languages including English, Chinese, German, Spanish, French, Italian, Japanese, and Korean. The model has 0.1B parameters and supports 44.1 kHz audio.
Produces zero-shot text-to-speech with voice cloning and emotion control in multiple languages. Requires 0.8B parameters and 6GB VRAM to output 22.05 kHz audio.
Produces Dutch text-to-speech audio from input text. Requires the transformers library, the microsoft/speecht5_hifigan vocoder, and a speaker embedding.
It converts text into spoken audio. The model has 2.2B parameters and uses the transformers library.
Produces 3D assets from a single image including baked normal and ambient occlusion maps. Runs on consumer hardware in ComfyUI without custom nodes or non-commercial dependencies.
Provides supplemental model checkpoints for research purposes. Model weights are available under the Apache-2.0 license.
The library provides a collection of WebGPU kernels for machine learning workloads. The collection contains 207 kernels and is licensed under Apache-2.0.
It provides a flashcard application for Android. It runs on Android devices.
7 releases71 clusters from 78 items
The model produces text output from text input. The model features 770B total parameters, 49B active parameters, a 1M token context window, and a size of 1.56TB.
The platform hosts open-weight AI models and benchmarks for developers. The source text contains no hardware or performance specifications.
The model produces reasoning outputs for a 262K context window. The model requires 102GB of total RAM and VRAM.
Produces text-to-speech audio in English, Portuguese, French, and German. Uses INT8 quantization for 20-second clips under an Apache-2.0 license.
The model produces text-to-speech audio in ten languages. The model uses 0.6B parameters and FP4 weights.
The model produces Japanese text-to-speech audio for the character Chi from the anime Chobits. It is a GPT-SoVITS v2Pro fine-tuned model trained on 483 audio samples.
Produces Khakas language speech from text input without requiring manual stress marks. Utilizes the Silero v5-cis base engine for processing.
7 releases61 clusters from 73 items
There's a lot of capital pouring into the business of giving models away. No specs stated in the source — open the link before quoting numbers.
The model produces multimodal reasoning. It requires 102GB of total memory.
The model produces multimodal outputs using a Mixture of Experts architecture. The model features 125B parameters and is available in GGUF format.
It produces speech for 10 languages from text input. The model has 0.6B parameters and uses FP4 weights.
Produces speech and audio from text. No hardware requirements are listed.
Produces Japanese text-to-speech audio for the character Chi from the anime Chobits. This is a fine-tuned GPT-SoVITS v2Pro model based on 483 audio clips.
Produces text-to-speech audio and voice. Requires GPT-SoVITS v2 and matching GPT and SoVITS weight files.
9 releases202 clusters from 217 items
The model produces text using gated residual, sparse attention, and per-layer embedding. It requires hardware to support 198 billion parameters.
The model produces text. The model has 27B parameters.
Produces multimodal reasoning outputs. 125B parameter model requiring 102GB total memory.
This model produces multimodal output. It features 125B parameters with 6B active and is available in GGUF format.
Ox Alpha produces code, handles sustained agentic work, and manages production workloads with complex reasoning and visual context. The model is released with open weights.
Granite 4.2 is a decoder-only language model with a 128,000-token context window. The model is available in 3B, 8B, and 30B parameter variants.
Provides bug fixes, auto compaction, LAN access, and improved streaming performance. The model has 27B parameters.
The model generates text-to-speech audio in English and Chinese. The model runs using the transformers library with PyTorch and CUDA support.
It provides retrainable robot behaviors such as walking, sitting, and grabbing trained in physics simulations for deployment on hardware. The project is licensed under Apache-2.0.
13 releases208 clusters from 218 items
Produces text and multimodal content. Requires hardware capable of supporting 198 billion parameters.
Identifies whether an AI agent produces false information or conceals known facts. Runs on a 27B parameter model.
This multimodal MoE model serves as a preview of the architecture used in Qwen4. The model has 125B parameters in GGUF format.
The model produces reasoning for coding, sustained agentic work, and workflows combining text with visual context. It runs on released open weights.
This decoder-only text model features a 128,000-token context window and supports tool use. The model is available in 3B, 8B, and 30B parameter variants.
The update provides bug fixes, auto compaction for long conversations, LAN remote access, and Qwen3.8-27B GGUF models. The 27B parameter model runs on hardware supporting MLX, RDNA GPUs, and custom llama.cpp builds.
developerjeremylive published SenseNova-U1.5-8B-MoT-etheroi to the Hugging Face hub for any-to-any. Specs: 8B params · Apache-2.0 · 4K.
Produces text-to-speech voice and audio. Runs on an RK3576 board via a Docker container.
The model produces text-to-speech audio and voice cloning. The model contains 4B parameters and utilizes the transformers library.
It produces speech and audio from text with support for voice cloning and design in English and Chinese. The model runs using the transformers and pytorch libraries on CUDA hardware.
The system provides a virtual executive team for strategy, finance, HR, legal, operations, marketing, and product roadmap. No hardware or software requirements are listed in the source.
It provides point-to-point WireGuard-encrypted tunnels between machines using a token-based exchange. It runs as a command-line tool or as an importable Go library.
Provides offline map access for navigation in areas without cellular signal. Licensed for non-commercial use.
8 releases172 clusters from 176 items
The update adds bug fixes for MLX and AMD hardware, LAN access, auto-compaction, and custom llama.cpp build support. The model utilizes 27B parameters.
The model provides any-to-any multimodal capabilities including text, image generation, and image editing. The model has 8B parameters and uses the transformers library.
The model produces 48kHz audio in 30 languages using a tokenizer-free, diffusion autoregressive architecture. The model has 2B parameters and is released under an Apache-2.0 license.
The model produces text-to-speech audio and zero-shot voice cloning in multiple languages. The model uses 0.1B parameters and a 44.1 kHz sampling rate.
The model produces zero-shot text-to-speech audio in multiple languages. The model has 0.1 billion parameters and a 44.1 kHz sample rate.
Produces text-to-speech audio for the Yoruba language. Requires gated access to download VITS weights in safetensors format.
Converts text into voice outputs for Chinese, Japanese, and English. Operates on the GPT-SoVITS library with a 32 kHz sample rate.
The model produces speech and audio from text input. The model has 8B parameters.
5 releases70 clusters from 71 items
The model performs any-to-any tasks across 140 languages with a 128K context window. The model has 26 billion total parameters and 4 billion active parameters.
Produces Amharic text using four 50M-parameter model architectures. Requires the transformers library and follows an Apache-2.0 license.
Produces Japanese text-to-speech audio using the GPT-SoVITS V2 model. Requires the GPT-SoVITS repository and weight files of 155 MB and 75 MB.
The model produces speech from text and handles image-to-text and text-to-image tasks. The model requires the pytorch library.
The model performs any-to-any tasks. The model runs using the diffusers library and safetensors weights.
12 releases180 clusters from 187 items
- Faster inference: up to 3.18 throughput improvement on a GPU and up to 2.87x on-device. Specs: 8B total / 1B active · 2.6B params · GGUF.
The update provides auto-compaction for conversations exceeding context limits, LAN access, and improved streaming performance. It runs on 27B parameter models in GGUF format.
Mojo🔥 is now open source The Mojo programming language has been promising an open source release since May 2023 . Specs: 27B params · Apache-2.0.
The model produces speech in 17 languages and clones voices from 6-second audio clips. The model features 7B parameters and produces 24 kHz audio.
Produces speech from text and performs image-to-text and text-to-image tasks. Requires PyTorch and external source code from a separate GitHub repository for inference.
The model produces text-to-speech audio and instant voice cloning. The system requires OpenVoice V1 and V2 base speaker and tone-color-converter checkpoints.
Produces Chinese text-to-speech with independent speaker and accent references. Uses a VITS-style acoustic model and HiFiGAN vocoder at 16 kHz.
It produces text-to-speech audio with controllable frame rates. It has 5B parameters and runs with the transformers library.
The model produces speech audio from text input. The model has 7B parameters.
The model produces text-to-speech audio for Chinese, English, German, Spanish, French, Italian, Japanese, and Korean. The model has 0.1B parameters and uses the transformers library.
Produces text-to-speech voice and audio. Operates at 4 kHz with a single-step time of 181 ms and a real-time factor of 0.3–1.1.
The model produces video with integrated stereo audio, dialogue, and foley in one pass. The model runs on local hardware or Comfy Cloud.
7 releases170 clusters from 187 items
The tool measures average and extreme fidelity loss for vendor-hosted LLM APIs using repeated-request and long-sequence analyses. The audit functions as a black-box system requiring no probability information from the target API.
It is a programming language optimized for GPU programming. The compiler and toolchain are released under an Apache-2.0 license.
The model functions as a vision-capable large language model. The model has 27B parameters and is available as a 17GB Q4_K_M quantized GGUF build.
The Qwen 3.8 2.4T model produces a Call of Duty clone from a single prompt. The model requires a B200 cluster to run.
The model produces text-to-speech for Amharic and Geez. The model supports FP8.
Produces Mandarin text-to-speech audio using a diffusion-based acoustic model. Requires the PyTorch library for execution.
The system allows users to build, edit, and run ComfyUI workflows via plain language and connects local and cloud environments to MCP clients. The system requires a local ComfyUI installation, a Comfy account, and an MCP-compatible client such as Claude, Cursor, or Codex.
5 releases123 clusters from 136 items
The model produces coordinate-aware token-space representations to reconstruct spatio-temporal queries from video data. It requires one feed-forward pass and a shared decoder.
It produces Arabic-centric text and performs on culturally grounded benchmarks. The model comes in 70B and 8B parameter versions.
Produces text and vision-capable outputs. Runs on a 128GB M5 Max MacBook Pro or NVIDIA DGX Spark using a 17GB Q4_K_M quantized build.
The model performs named entity recognition in Amharic. It requires the transformers library and Hugging Face model weights.
Produces text-to-speech audio in English and several Indian languages. No hardware requirements are provided.
3 releases65 clusters from 68 items
The two leading labs are already feeling extremely threatened by Qwen & friends, however I think there are tons of enterprises and organizations in the West that don't feel comfortable using Chinese models. Specs: 120B params.
Models and datasets on HF hub are growing on a daily basis. Specs: 754B params.
The system records robot demonstrations, trains policies on the data from Hugging Face Hub, and deploys them to hardware. Licensed under Apache-2.0.
9 releases180 clusters from 191 items
The model produces text responses through an LLM interface. The weights are available for 1.7 trillion parameters and require 893 GB of storage.
Produces text output via a mixture-of-experts architecture for agentic tasks. Requires 30B total parameters and 3B active parameters.
Produces text outputs based on text, image, video, and audio inputs. Mixture-of-Experts model with 280B total parameters, 16B activated parameters, and a 512K token context length.
DeepSeek Harness is an agent framework where components like models, tools, and sandboxes are implemented as plugins. The code is released under the MIT license.
Generates full songs up to five minutes in length based on lyrics and a description of the sound. The model contains 8B parameters and produces 32 kHz stereo audio.
Highlights 582 PRs from 194 contributors. Specs: 1M context · 8K · FP8.
The system manages recording robot demonstrations, training policies on them using LeRobot data formats, and deploying the models to hardware. This software is licensed under Apache-2.0.
Produces embedding vectors for earth observation data analysis. Available in 8-bit format.
The model powers coding agent applications and long-running personal assistants. It runs on all platforms including Apple Silicon, NVIDIA, and AMD hardware.
8 releases159 clusters from 173 items
Produces text for agentic use cases, coding, document analysis, and personal assistants. 30B parameter dense model with a 2B ViT-style vision encoder and 28B text decoder under Apache-2.0 license.
The system identifies differences between benchmark scores and deployment performance by masking primary candidate tokens at word boundaries during runtime. No specific technical specifications are provided in the source text.
Produces text outputs from a mixture-of-experts model for agent workflows. Requires 30B total parameters with 3B active parameters.
Produces text and long context outputs. Requires 104B parameters, up to 162TB for lossless inference, and compatible with GGUF format.
The model generates speech for multiple languages. It supports 12 languages and requires open weights deployment.
The model generates multi-shot videos with synchronized audio and frames at a rate of 50 fps. It uses 12B parameters and runs on local GPUs to output video in 720p or 4K resolutions.
Generates video from images and text. Requires access to the Hugging Face repository via a gated license.
Kimi K3 produces multimodal content with a 1M-token context and MoonViT3d vision tower. The model uses 2.8T parameters, FP8 quantization, and supports SGLang serving with DCP and speculative decoding.
8 releases114 clusters from 118 items
Provides a graphical interface for stable diffusion and other image generation models. Requires PyTorch 2.7 or higher to run.
Unsloth Desktop is here! Specs: 30B params · GGUF.
The model generates text using a mixture-of-experts architecture with 3B active parameters. It requires an infrastructure supporting the NVIDIA NemoClaw stack and is available via Ollama.
The model produces end-to-end agentic task completion, tool use, and multi-step reasoning. The model has 30 billion parameters and is under an Apache-2.0 license.
Muse Glimmer produces text outputs for large language model tasks. The model has 30 billion parameters and a 128,000-token context window.
This system enables llama.cpp to utilize advanced Metal kernels for inference on macOS virtual machines. Requires a host with Apple Silicon and the Lume virtualization stack.
SGLang supports Kimi K3 and MiniMax-H3 models for multimodal, video, and audio generation. The system supports 1M context, 8K resolution, and FP8 precision.
Muse Glimmer produces coding agent outputs and responses for personal assistant frameworks. Requires Ollama software and runs on NVIDIA, AMD, and Apple Silicon hardware.
3 releases116 clusters from 118 items
To celebrate, we are shipping with Meta day-0 support in transformers, llama.cpp, vLLM, Inference Endpoints, and other libraries. Specs: 30B params.
The model produces text and processes vision input for coding, document analysis, and personal assistant tasks. It is a 30B parameter dense model with a 2B ViT-style encoder and a 28B parameter text decoder.
Muse Glimmer is a multimodal model for agent workloads and coding applications. The model has 30B parameters.
1 release23 clusters from 24 items
WorldClaw generates 3D content. No specific hardware or software requirements are listed.
1 release114 clusters from 116 items
It identifies steerable semantic concept directions to control behavior in multiple independently trained large language models. The study tests five open-weight models with parameter counts from 0.8B to 8B.
3 releases133 clusters from 136 items
Produces multilingual text generation across ten languages with a 256K token context length. Operates as a Mixture-of-Experts model with 750B total parameters and 37B activated per token.
Ollama provides a local runner for large language models with updated TUI features and support for the Qwen3.5 model. The software runs on Ollama v0.32.6.
Produces video clips of 20 seconds. Requires ComfyUI and Partner Nodes.