Pestle-27B-Ternary

Pestle-27B-Ternary

Pestle-27B-Ternary is a compact 27B-class ternary-quantized Qwen3.6-derived model packaged as a single runnable GGUF for local inference, optimized for medical and general-purpose text generation, biomedical QA, and pharmaceutical retrieval (research preview).

Pestle-27B-Ternary is healthcare software teams evaluate for healthcare. Use this page to review pricing, integration signals, and the best alternatives before you commit.

Contact for pricing
One of 42 tools in Healthcare
Added 1 month ago
Data reviewed Aug 15, 2026

Profile facts come from the vendor source. AiMatch labels unknown pricing or API details instead of estimating them.

Review official source →

Quick Overview

Best for: Healthcare

What it does

Healthcare software for decision-makers comparing workflow fit and alternatives.

Best fit

Healthcare

Pricing snapshot

Contact for pricing

Next step

Compare Pestle-27B-Ternary with similar tools before you shortlist it.

Compare this tool before you shortlist it

Review alternatives, pricing posture, and workflow fit side by side.

Pestle-27B-Ternary is a ternary-compressed variant of a Qwen3.6-derived 27B model packaged as a single 8.48 GB GGUF for local inference. It is released as a research preview by Doses AI for evaluation and development of medical-text workflows, biomedical question answering, pharmaceutical retrieval, coding, and general assistant tasks. The model emphasizes compact runnable size, strong medical-bench performance, and local execution on Apple Silicon (Metal), NVIDIA GPUs (CUDA), or CPU fallback via Mortar/llama.cpp-compatible runtimes. Pestle is explicitly not a medical device and is not intended for clinical decision-making; outputs require professional review and validation for clinical use.

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Own this listing?

Claim this page for a one-time $29 to add pricing, features, screenshots, verified owner details, and a clearly labeled 30-day category position after the profile is live.

Claim this listing for $29

Key Features

Single-file runnable GGUF

Distributed as one 8.48 GB GGUF artifact (pestle-27b-ternary.gguf) that contains the runnable ternary model and embedded chat template, with no separate overlay or conversion required.

Ternary compression with BF16 decoder block

Applies Doses AI ternary compression with a matching-parent BF16 final decoder block (hybrid precision) to reduce footprint while retaining a 27B-class capability.

Strong medical and biomedical performance

Benchmarked across medical datasets (MedQA, PubMedQA, BioASQ, PharmaRAG, MMLU medical subjects) with reported scores presented on the model page.

Multi-hardware local inference

Runs locally with Mortar (selects Metal on macOS, CUDA on Linux when NVIDIA GPU available, or CPU fallback). Also supports llama.cpp, vLLM, Ollama, Docker Model Runner, and other local runtimes.

Optional vision input

Supports optional image/vision input via a separately downloadable mmproj projection file (mmproj-pestle-27b-ternary.gguf).

Very long context

Advertised support for context lengths up to 262K tokens (practical limits depend on memory).

Apache 2.0 license

Released under the Apache-2.0 license.

Pricing

Free Tier Available

Model artifact is available for download on Hugging Face under the Apache-2.0 license (no pricing listed).

Use Cases

Medical QA and clinical knowledge evaluation

Use for research evaluation and development of medical question-answering systems and clinical knowledge retrieval (benchmarked on MedQA, MedMCQA, MMLU clinical subjects).

Biomedical retrieval and pharmaceutical RAG

Designed for pharmaceutical retrieval and retrieval-augmented generation workflows (PharmaRAG benchmarks reported).

On-device/local inference for privacy

Run locally on-device or on-premises to support data-residency and privacy-sensitive workflows while retaining model capability.

Coding and general assistant tasks

Supports coding tasks and general assistant interactions; coding benchmarks (HumanEval+, MBPP+) are reported.

Research & development

Intended as a research preview for evaluation, benchmarking, and development of locally hosted systems rather than production clinical deployment.

Integrations

llama.cpp

Run local OpenAI-compatible server (llama serve) or CLI (llama cli) against the GGUF; pre-built binaries and releases referenced.

Mortar (Doses AI runtime)

Mortar runtime is recommended for optimized local inference; build instructions and runtime selection (Metal/CUDA/CPU) are provided.

vLLM

vLLM can serve the model via an OpenAI-compatible API (pip install vllm; vllm serve "Doses-AI/Pestle-27B-Ternary-GGUF").

Ollama

Model can be run with Ollama using the hf.co model identifier (ollama run hf.co/Doses-AI/Pestle-27B-Ternary-GGUF).

Docker Model Runner / Docker

Docker model runner commands are provided (docker model run hf.co/Doses-AI/Pestle-27B-Ternary-GGUF) for containerized execution.

Unsloth Studio, Lemonade, Pi, OpenClaw, Hermes Agent, LM Studio, Jan, vLLM-compatible tooling

Multiple local UIs and agents are supported or have usage instructions in the model card for interactive chat and integration.

Benefits

Compact single-file GGUF reduces deployment complexity and enables local inference from a single runnable artifact.
Ternary compression and hybrid precision lower storage and memory costs while maintaining 27B-class capability.
Local execution on Metal, CUDA, or CPU supports privacy and on-premises data-residency requirements.
Benchmarked medical and biomedical performance provides evidence for research and evaluation use cases.

Limitations

Research preview: not a medical device and not intended for diagnosis, treatment, prescribing, triage, or patient-management decisions.
Outputs may be inaccurate, incomplete, biased, or confidently wrong; medical outputs require review by qualified professionals.
Local CPU-only inference is substantially slower and requires sufficient system RAM for model and context.
Long outputs can occasionally become repetitive; production systems should use generation limits and monitoring.
Local execution supports privacy goals but does not by itself establish regulatory compliance; do not send identifiable patient information to unapproved environments.

Frequently Asked Questions

No verified FAQs are available.

Getting Started

  1. 1 Step 1: Download the model GGUF from Hugging Face (hf download Doses-AI/Pestle-27B-Ternary-GGUF pestle-27b-ternary.gguf --local-dir models/Pestle-27B-Ternary).
  2. 2 Step 2: Build or install a compatible runtime (e.g., clone and build Mortar: git clone https://github.com/DosesAI/mortar.cpp.git && ./scripts/build-mortar.sh) or install llama.cpp/vLLM/Ollama as described on the model page.
  3. 3 Step 3: Run the model locally (example: ./mortar --model ../models/Pestle-27B-Ternary/pestle-27b-ternary.gguf) or start a local OpenAI-compatible server using llama serve or vllm serve and call via the OpenAI-compatible API.

Support

docs

Model card and instructions on the Hugging Face model page (download, runtimes, recommended settings).

github

Mortar runtime and other referenced repositories (e.g., https://github.com/DosesAI/mortar.cpp.git) provide build instructions and issue trackers for runtime support.

community

Hugging Face community, discussion on the model page, and upstream project repositories for questions and troubleshooting.

API

Available: No

Compare Pestle-27B-Ternary with similar tools

See how it stacks up against alternatives

Related Tools

View all 42 β†’
Free
Kith

Kith

Kith is an AI-driven clinical workspace for therapists and clinical psychologists that provides ambient transcription, AI-written SOAP notes, scheduling, and session recording to reduce paperwork and streamline practice management.

Healthcare
Contact for pricing
Codebaby

Codebaby

CodeBaby builds real-time interactive AI-driven digital humans and avatars for businesses and organizations, offering customizable avatars, turnkey managed solutions, and hologram-enabled deployments to improve engagement across education, healthcare, hospitality, retail, events, and customer service.

Healthcare
Contact for pricing
mpathic

mpathic

mpathic is a human-centered AI safety company that helps AI builders evaluate, stress-test, and improve human-facing models using expert-led red teaming, clinically grounded benchmarking, and tools like mpathic Studio and an API.

Healthcare
Enterprise-ready
Contact for pricing
morphik

morphik

Morphik provides AI workers that automate accounts payable, billing, collections, care management, and other back-office workflows for skilled nursing and senior living operators to reduce manual work and increase NOI.

Healthcare
Enterprise-ready
Hubble

Hubble

Hubble provides AI-native, patient-mediated medical record retrieval that assembles a single, source-traceable health record by connecting EHRs, payers, HIE networks and using voice/browser agents for hard-to-reach sources.

Healthcare
Free
Sleepi

Sleepi

Sleepi is an AI-driven sleep coach app that combines artificial intelligence and sleep science to deliver personalized, emotionally aware sleep coaching, optimize daily energy and performance, and provide habit-based lifestyle recommendations.

Healthcare
Sonde Health

Sonde Health

Sonde Health provides a clinically validated vocal biomarker platform that uses voice AI and machine learning to analyze everyday speech and deliver continuous insights into mental, cognitive, and respiratory health via an API/SDK for integration into apps and devices.

Healthcare
woebot-health

woebot-health

Woebot Health develops chat-based AI wellness tools intended to make mental health support more accessible. The company describes its approach as combining empathy and clinical rigor and says it works with both individuals and organizations.

Healthcare

Premium Alternatives

Paid
Revmaxx

Revmaxx

RevMaxx provides AI-powered healthcare operations software including an ambient AI medical scribe with deep EHR integration, agentic RCM automation, and a telehealth peptide therapy platform to streamline documentation, billing, and telehealth workflows for medical practices and billing organizations.

Healthcare
Enterprise-ready

Explore Related Categories