Conway Research releases a 7.89 GB AI model tuned for tool calling
Saluki’s model card reports higher tool-calling scores than the full-size baseline and 96% average benchmark retention. Those figures are Conway’s own comparisons, not an independent ranking.
Listen to this story
The audio brief
Story brief
3 key pointsUnderdog Saluki 27B 1.0 packages a 2-bit version of Qwen3.8-27B as a 7.89 GB GGUF file, with support for stock llama.cpp rather than a Conway-specific runtime. That makes local deployment more accessible for users prioritizing tool use, though the headline size excludes an optional vision add-on. Conway reports 96% average performance retention across nine benchmarks; its tool-calling gains are company-reported results, not evidence that the compressed model outperforms its parent on every task.
- 01
Conway reports Underdog Bench tool-calling scores of 88 for Saluki and 84 for the 54 GB full model; parallel tool-call scores are 42 and 35.
- 02
The optional vision files add 928 MB in F16 or 629 MB in Q8_0, beyond the 7.89 GB language-model file.
- 03
The repository lists an Apache 2.0 license, and the standard GGUF packaging supports deployment through llama.cpp and compatible apps.
Conway Research has released a compressed version of Qwen3.8-27B whose main model file is 7.89 GB, compared with 54 GB for the full model. Called Underdog Saluki 27B 1.0, the download is tuned to preserve tool calling. Conway’s own benchmark comparisons show higher scores for that capability despite the much smaller package.
A compressed Qwen model in a standard file
Saluki is a quantized model—a compressed version of an existing model—rather than a separate base model. Conway identifies Qwen3.8-27B as its starting point and labels the release as 2-bit. The downloadable language-model file is named Underdog-Saluki-27B-1.0-IQ2-mix.gguf.
The packaging is meant to work with stock llama.cpp, software for running models locally, and apps built on it. Conway describes the file as a standard GGUF, the model-file format used here. That gives the release a concrete deployment route: download the model and use the existing runtime, rather than a Conway-specific version of it.
Tool calling is the stated priority
Conway says Saluki is tuned to keep tool calling intact: the ability to request functions or tools rather than only return conversational text. Its headline comparison is Underdog Bench, where the model card reports a tool-calling score of 88 for Saluki versus 84 for the full model.
The parallel tool-call comparison also favors the compressed release, at 42 versus 35. Separately, Conway reports 96% average performance retention across nine benchmarks. That figure describes retained benchmark performance relative to the parent model; it is not a claim that Saluki answers 96% of all requests correctly.
These are Conway’s reported results. The two tool-calling comparisons support its stated focus, but they should not be read as a finding that the compressed model is better across every task. The average-retention figure makes a different claim: much of the parent model’s measured performance remains, not that every benchmark improves.
Image input adds a separate file
The main package handles text. Conway also offers an optional vision add-on for image input, so the 7.89 GB headline size is not the complete download for someone choosing that capability. The model card lists two versions of the additional file:
- F16 vision add-on: 928 MB.
- Q8_0 vision add-on: 629 MB.
The repository lists an Apache 2.0 license. Together with the standard runtime support and separate vision files, that makes Saluki a downloadable model release with explicit packaging choices—not merely a benchmark result or access to a hosted assistant.
Sources
- huggingface.coConwayResearch/Underdog-Saluki-27B-1.0 · Hugging Face
Reader comments
Newest comments first. Replies stay oldest first.