Pestle-27B-Ternary

Pestle-27B-Ternary

Pestle-27B-Ternary is a compact 27B-class ternary-quantized Qwen3.6-derived model packaged as a single runnable GGUF for local inference, optimized for medical and general-purpose text generation, biomedical QA, and pharmaceutical retrieval (research preview).

Pestle-27B-Ternary is healthcare software teams evaluate for healthcare. Use this page to review pricing, integration signals, and the best alternatives before you commit.

Contact for pricing
#26 in Healthcare (26 tools)
Just launched
Data reviewed Aug 15, 2026

Profile facts come from the vendor source. AiMatch labels unknown pricing or API details instead of estimating them.

Review official source →

Quick Overview

Best for: Healthcare

What it does

Healthcare software for decision-makers comparing workflow fit and alternatives.

Best fit

Healthcare

Pricing snapshot

Contact for pricing

Next step

Compare Pestle-27B-Ternary with similar tools before you shortlist it.

Compare this tool before you shortlist it

Review alternatives, pricing posture, and workflow fit side by side.

Pestle-27B-Ternary

Pestle-27B-Ternary is a ternary-compressed variant of a Qwen3.6-derived 27B model packaged as a single 8.48 GB GGUF for local inference. It is released as a research preview by Doses AI for evaluation and development of medical-text workflows, biomedical question answering, pharmaceutical retrieval, coding, and general assistant tasks. The model emphasizes compact runnable size, strong medical-bench performance, and local execution on Apple Silicon (Metal), NVIDIA GPUs (CUDA), or CPU fallback via Mortar/llama.cpp-compatible runtimes. Pestle is explicitly not a medical device and is not intended for clinical decision-making; outputs require professional review and validation for clinical use.

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Own this listing?

Claim this page for a one-time $29 to add pricing, features, screenshots, verified owner details, and a clearly labeled 30-day category position after the profile is live.

Claim this listing for $29

Key Features

Single-file runnable GGUF

Distributed as one 8.48 GB GGUF artifact (pestle-27b-ternary.gguf) that contains the runnable ternary model and embedded chat template, with no separate overlay or conversion required.

Ternary compression with BF16 decoder block

Applies Doses AI ternary compression with a matching-parent BF16 final decoder block (hybrid precision) to reduce footprint while retaining a 27B-class capability.

Strong medical and biomedical performance

Benchmarked across medical datasets (MedQA, PubMedQA, BioASQ, PharmaRAG, MMLU medical subjects) with reported scores presented on the model page.

Multi-hardware local inference

Runs locally with Mortar (selects Metal on macOS, CUDA on Linux when NVIDIA GPU available, or CPU fallback). Also supports llama.cpp, vLLM, Ollama, Docker Model Runner, and other local runtimes.

Optional vision input

Supports optional image/vision input via a separately downloadable mmproj projection file (mmproj-pestle-27b-ternary.gguf).

Very long context

Advertised support for context lengths up to 262K tokens (practical limits depend on memory).

Apache 2.0 license

Released under the Apache-2.0 license.

Pricing

Free Tier Available

Model artifact is available for download on Hugging Face under the Apache-2.0 license (no pricing listed).

Use Cases

Medical QA and clinical knowledge evaluation

Use for research evaluation and development of medical question-answering systems and clinical knowledge retrieval (benchmarked on MedQA, MedMCQA, MMLU clinical subjects).

Biomedical retrieval and pharmaceutical RAG

Designed for pharmaceutical retrieval and retrieval-augmented generation workflows (PharmaRAG benchmarks reported).

On-device/local inference for privacy

Run locally on-device or on-premises to support data-residency and privacy-sensitive workflows while retaining model capability.

Coding and general assistant tasks

Supports coding tasks and general assistant interactions; coding benchmarks (HumanEval+, MBPP+) are reported.

Research & development

Intended as a research preview for evaluation, benchmarking, and development of locally hosted systems rather than production clinical deployment.

Integrations

llama.cpp

Run local OpenAI-compatible server (llama serve) or CLI (llama cli) against the GGUF; pre-built binaries and releases referenced.

Mortar (Doses AI runtime)

Mortar runtime is recommended for optimized local inference; build instructions and runtime selection (Metal/CUDA/CPU) are provided.

vLLM

vLLM can serve the model via an OpenAI-compatible API (pip install vllm; vllm serve "Doses-AI/Pestle-27B-Ternary-GGUF").

Ollama

Model can be run with Ollama using the hf.co model identifier (ollama run hf.co/Doses-AI/Pestle-27B-Ternary-GGUF).

Docker Model Runner / Docker

Docker model runner commands are provided (docker model run hf.co/Doses-AI/Pestle-27B-Ternary-GGUF) for containerized execution.

Unsloth Studio, Lemonade, Pi, OpenClaw, Hermes Agent, LM Studio, Jan, vLLM-compatible tooling

Multiple local UIs and agents are supported or have usage instructions in the model card for interactive chat and integration.

Benefits

Compact single-file GGUF reduces deployment complexity and enables local inference from a single runnable artifact.
Ternary compression and hybrid precision lower storage and memory costs while maintaining 27B-class capability.
Local execution on Metal, CUDA, or CPU supports privacy and on-premises data-residency requirements.
Benchmarked medical and biomedical performance provides evidence for research and evaluation use cases.

Limitations

Research preview: not a medical device and not intended for diagnosis, treatment, prescribing, triage, or patient-management decisions.
Outputs may be inaccurate, incomplete, biased, or confidently wrong; medical outputs require review by qualified professionals.
Local CPU-only inference is substantially slower and requires sufficient system RAM for model and context.
Long outputs can occasionally become repetitive; production systems should use generation limits and monitoring.
Local execution supports privacy goals but does not by itself establish regulatory compliance; do not send identifiable patient information to unapproved environments.

Frequently Asked Questions

Claim this listing to publish FAQs.

Getting Started

  1. 1 Step 1: Download the model GGUF from Hugging Face (hf download Doses-AI/Pestle-27B-Ternary-GGUF pestle-27b-ternary.gguf --local-dir models/Pestle-27B-Ternary).
  2. 2 Step 2: Build or install a compatible runtime (e.g., clone and build Mortar: git clone https://github.com/DosesAI/mortar.cpp.git && ./scripts/build-mortar.sh) or install llama.cpp/vLLM/Ollama as described on the model page.
  3. 3 Step 3: Run the model locally (example: ./mortar --model ../models/Pestle-27B-Ternary/pestle-27b-ternary.gguf) or start a local OpenAI-compatible server using llama serve or vllm serve and call via the OpenAI-compatible API.

Support

docs

Model card and instructions on the Hugging Face model page (download, runtimes, recommended settings).

github

Mortar runtime and other referenced repositories (e.g., https://github.com/DosesAI/mortar.cpp.git) provide build instructions and issue trackers for runtime support.

community

Hugging Face community, discussion on the model page, and upstream project repositories for questions and troubleshooting.

API

Available: No

Compare Pestle-27B-Ternary with similar tools

See how it stacks up against alternatives

Related Tools

View all 26 →
Free
Fluidforms

Fluidforms

FluidForms is an intelligent intake platform that uses adaptive AI, document intelligence, and dynamic conversational forms to capture complete, structured data for intakes such as legal claims, medical records, hiring, customer feedback, and more.

Healthcare
Freemium
Medankigen

Medankigen

MedAnkiGen is an AI-powered Anki flashcard generator that converts PDFs, PowerPoints, Word documents, and notes into high-yield Anki decks (including cloze deletions and image occlusion) and exports ready-to-import .apkg files. It is built for medical exams (USMLE, NCLEX, MCAT) but works for any subject, and offers a free plan to get started.

Healthcare
Freemium
Foodintake

Foodintake

FoodIntake is a mobile app for effortless meal logging and nutrition analysis that uses AI to identify foods, estimate calories and provide detailed nutrient breakdowns to support weight and dietary goals.

Healthcare
Free
medical-chat

medical-chat

Medical Chat is a HIPAA-ready clinical AI assistant that answers clinician questions in natural language, provides differential diagnoses, dosing guidance, structured diagnosis reports, and evidence-grounded citations from PubMed, UpToDate, NICE and major guidelines.

Healthcare
High-growth
Contact for pricing
Sunoh

Sunoh

Sunoh.ai is an AI-powered medical scribe that listens to patient-provider conversations, generates clinical notes in seconds, and integrates with major EHRs to reduce documentation time and provider burnout.

Healthcare
Free
Easygpt

Easygpt

Inbound GPT by EasyGPT Builders is a no-code, HIPAA-compliant builder for human-in-the-loop conversational agents that handle phone calls, chat, SMS, email, and video—designed primarily for healthcare practices to automate patient communication and scheduling while handing off to humans as needed.

Healthcare
Free
Sleepi

Sleepi

Sleepi is an AI-driven sleep coach app that combines artificial intelligence and sleep science to deliver personalized, emotionally aware sleep coaching, optimize daily energy and performance, and provide habit-based lifestyle recommendations.

Healthcare
Paid
Revmaxx

Revmaxx

RevMaxx provides AI-powered healthcare operations software including an ambient AI medical scribe with deep EHR integration, agentic RCM automation, and a telehealth peptide therapy platform to streamline documentation, billing, and telehealth workflows for medical practices and billing organizations.

Healthcare
Enterprise-ready

Premium Alternatives

Paid
Revmaxx

Revmaxx

RevMaxx provides AI-powered healthcare operations software including an ambient AI medical scribe with deep EHR integration, agentic RCM automation, and a telehealth peptide therapy platform to streamline documentation, billing, and telehealth workflows for medical practices and billing organizations.

Healthcare
Enterprise-ready

Explore Related Categories