Pestle-27B-Ternary
Pestle-27B-Ternary is a compact 27B-class ternary-quantized Qwen3.6-derived model packaged as a single runnable GGUF for local inference, optimized for medical and general-purpose text generation, biomedical QA, and pharmaceutical retrieval (research preview).
Pestle-27B-Ternary is healthcare software teams evaluate for healthcare. Use this page to review pricing, integration signals, and the best alternatives before you commit.
Profile facts come from the vendor source. AiMatch labels unknown pricing or API details instead of estimating them.
Review official source →Quick Overview
Best for: Healthcare
What it does
Healthcare software for decision-makers comparing workflow fit and alternatives.
Best fit
Healthcare
Pricing snapshot
Contact for pricing
Next step
Compare Pestle-27B-Ternary with similar tools before you shortlist it.
Compare this tool before you shortlist it
Review alternatives, pricing posture, and workflow fit side by side.
Pestle-27B-Ternary
Pestle-27B-Ternary is a ternary-compressed variant of a Qwen3.6-derived 27B model packaged as a single 8.48 GB GGUF for local inference. It is released as a research preview by Doses AI for evaluation and development of medical-text workflows, biomedical question answering, pharmaceutical retrieval, coding, and general assistant tasks. The model emphasizes compact runnable size, strong medical-bench performance, and local execution on Apple Silicon (Metal), NVIDIA GPUs (CUDA), or CPU fallback via Mortar/llama.cpp-compatible runtimes. Pestle is explicitly not a medical device and is not intended for clinical decision-making; outputs require professional review and validation for clinical use.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Own this listing?
Claim this page for a one-time $29 to add pricing, features, screenshots, verified owner details, and a clearly labeled 30-day category position after the profile is live.
Claim this listing for $29Key Features
Single-file runnable GGUF
Distributed as one 8.48 GB GGUF artifact (pestle-27b-ternary.gguf) that contains the runnable ternary model and embedded chat template, with no separate overlay or conversion required.
Ternary compression with BF16 decoder block
Applies Doses AI ternary compression with a matching-parent BF16 final decoder block (hybrid precision) to reduce footprint while retaining a 27B-class capability.
Strong medical and biomedical performance
Benchmarked across medical datasets (MedQA, PubMedQA, BioASQ, PharmaRAG, MMLU medical subjects) with reported scores presented on the model page.
Multi-hardware local inference
Runs locally with Mortar (selects Metal on macOS, CUDA on Linux when NVIDIA GPU available, or CPU fallback). Also supports llama.cpp, vLLM, Ollama, Docker Model Runner, and other local runtimes.
Optional vision input
Supports optional image/vision input via a separately downloadable mmproj projection file (mmproj-pestle-27b-ternary.gguf).
Very long context
Advertised support for context lengths up to 262K tokens (practical limits depend on memory).
Apache 2.0 license
Released under the Apache-2.0 license.
Pricing
Model artifact is available for download on Hugging Face under the Apache-2.0 license (no pricing listed).
Use Cases
Medical QA and clinical knowledge evaluation
Use for research evaluation and development of medical question-answering systems and clinical knowledge retrieval (benchmarked on MedQA, MedMCQA, MMLU clinical subjects).
Biomedical retrieval and pharmaceutical RAG
Designed for pharmaceutical retrieval and retrieval-augmented generation workflows (PharmaRAG benchmarks reported).
On-device/local inference for privacy
Run locally on-device or on-premises to support data-residency and privacy-sensitive workflows while retaining model capability.
Coding and general assistant tasks
Supports coding tasks and general assistant interactions; coding benchmarks (HumanEval+, MBPP+) are reported.
Research & development
Intended as a research preview for evaluation, benchmarking, and development of locally hosted systems rather than production clinical deployment.
Integrations
llama.cpp
Run local OpenAI-compatible server (llama serve) or CLI (llama cli) against the GGUF; pre-built binaries and releases referenced.
Mortar (Doses AI runtime)
Mortar runtime is recommended for optimized local inference; build instructions and runtime selection (Metal/CUDA/CPU) are provided.
vLLM
vLLM can serve the model via an OpenAI-compatible API (pip install vllm; vllm serve "Doses-AI/Pestle-27B-Ternary-GGUF").
Ollama
Model can be run with Ollama using the hf.co model identifier (ollama run hf.co/Doses-AI/Pestle-27B-Ternary-GGUF).
Docker Model Runner / Docker
Docker model runner commands are provided (docker model run hf.co/Doses-AI/Pestle-27B-Ternary-GGUF) for containerized execution.
Unsloth Studio, Lemonade, Pi, OpenClaw, Hermes Agent, LM Studio, Jan, vLLM-compatible tooling
Multiple local UIs and agents are supported or have usage instructions in the model card for interactive chat and integration.
Benefits
Limitations
Frequently Asked Questions
Claim this listing to publish FAQs.
Getting Started
- 1 Step 1: Download the model GGUF from Hugging Face (hf download Doses-AI/Pestle-27B-Ternary-GGUF pestle-27b-ternary.gguf --local-dir models/Pestle-27B-Ternary).
- 2 Step 2: Build or install a compatible runtime (e.g., clone and build Mortar: git clone https://github.com/DosesAI/mortar.cpp.git && ./scripts/build-mortar.sh) or install llama.cpp/vLLM/Ollama as described on the model page.
- 3 Step 3: Run the model locally (example: ./mortar --model ../models/Pestle-27B-Ternary/pestle-27b-ternary.gguf) or start a local OpenAI-compatible server using llama serve or vllm serve and call via the OpenAI-compatible API.
Support
docs
Model card and instructions on the Hugging Face model page (download, runtimes, recommended settings).
github
Mortar runtime and other referenced repositories (e.g., https://github.com/DosesAI/mortar.cpp.git) provide build instructions and issue trackers for runtime support.
community
Hugging Face community, discussion on the model page, and upstream project repositories for questions and troubleshooting.
API
Compare Pestle-27B-Ternary with similar tools
See how it stacks up against alternatives
Related Tools
View all 26 →
Fluidforms
FluidForms is an intelligent intake platform that uses adaptive AI, document intelligence, and dynamic conversational forms to capture complete, structured data for intakes such as legal claims, medical records, hiring, customer feedback, and more.
Medankigen
MedAnkiGen is an AI-powered Anki flashcard generator that converts PDFs, PowerPoints, Word documents, and notes into high-yield Anki decks (including cloze deletions and image occlusion) and exports ready-to-import .apkg files. It is built for medical exams (USMLE, NCLEX, MCAT) but works for any subject, and offers a free plan to get started.
Foodintake
FoodIntake is a mobile app for effortless meal logging and nutrition analysis that uses AI to identify foods, estimate calories and provide detailed nutrient breakdowns to support weight and dietary goals.
medical-chat
Medical Chat is a HIPAA-ready clinical AI assistant that answers clinician questions in natural language, provides differential diagnoses, dosing guidance, structured diagnosis reports, and evidence-grounded citations from PubMed, UpToDate, NICE and major guidelines.
Easygpt
Inbound GPT by EasyGPT Builders is a no-code, HIPAA-compliant builder for human-in-the-loop conversational agents that handle phone calls, chat, SMS, email, and video—designed primarily for healthcare practices to automate patient communication and scheduling while handing off to humans as needed.
Revmaxx
RevMaxx provides AI-powered healthcare operations software including an ambient AI medical scribe with deep EHR integration, agentic RCM automation, and a telehealth peptide therapy platform to streamline documentation, billing, and telehealth workflows for medical practices and billing organizations.
Premium Alternatives
Revmaxx
RevMaxx provides AI-powered healthcare operations software including an ambient AI medical scribe with deep EHR integration, agentic RCM automation, and a telehealth peptide therapy platform to streamline documentation, billing, and telehealth workflows for medical practices and billing organizations.