LLMs
LLMs
Large language models (LLMs), APIs, agents, and harnesses.
Models
Qwen3.8-27B
- Unsloth: Qwen3.8-27B
- GSQ-RCO: Qwen3.8-27B
- MiaAI-Lab: Qwen3.8-27B on one DGX Spark.
- MiaAI-Lab: Qwen3.8-27B for one 24GB GPU or one DGX Spark.
Qwen3.8-Flash-Next
- Unsloth: Qwen3.8-Flash-Next
- qwen38-flash-next-spark on one DGX Spark.
Deepseek v4 flash
- MiaAI Lab: DS4f on one DGX Spark.
- Entrpi: DS4f on spark on one DGX Spark.
Gemma4
Speech recognition and text-to-speech
Document processing
Biology
Runtime
- FreeToken : Optimized for MoE models.
- Unsloth desktop
- Lemonade for AMD GPUs.
- Lucebox
- Club 3090 : recipes for 3090/4090/5090 owners.
llama.cpp
- llama.cpp GitHub repo
- beellama.cpp :
llama.cppfork supporting KVarN KV cache format. - ik_llama.cpp :
llama.cppfork with new quants and improved performance for MoE models.
vLLM
- vllm GitHub repo
- vllm-radiance docker image for 2 R9700’s.
- spark-vllm-docker : vLLM docker images for DGX sparks.
SGLang
MCP servers
Model Context Protocol (MCP)
Agents and harnesses
Last updated on