Fine-Tuning as a Service
On-Premise LLM Fine-Tuning for Healthcare
On narrowly defined workflows, a smaller model tuned and evaluated against local requirements can outperform a larger general-purpose model. We train the adapters on hospital-controlled hardware and data.
How does GofarAI fine-tune a model without moving patient data?
GofarAI performs tuning on the hospital’s own hardware and data. The base model stays frozen while a smaller adapter is trained, evaluated, and handed over with documentation. The intended workflow keeps patient data and the resulting adapter inside the hospital’s environment.
Why we do the tuning
Tuning is not self-serve.
Self-serve tuning in a hospital is a liability minefield. Bad training data produces a degraded model, and there’s no one accountable for what went wrong. Compliance teams need a named human on the hook for what went into the model.
Every tuning engagement is performed directly by GofarAI’s founder. Controlled quality. Named accountability.
Technical Approach
LoRA and QLoRA adapters, on your GPUs.
Freeze the base. Train an adapter.
The base model weights stay frozen. We train a small adapter (less than 1% of model size) that carries all the workflow-specific behavior. Adapters are portable, auditable, and cheap to iterate.
Fits the hardware you already have.
QLoRA fine-tunes a 7–8B model on a single 24GB GPU. A 13–14B model fits on one 48GB card or two 24GB cards. We reuse the same inference hardware for training — overnight or on weekends.
Training stays inside the hospital environment.
Training is designed to occur on your GPUs inside your environment without transferring training data to an external model provider. Your legal and compliance teams determine which agreements are required for the engagement. Our lab-side GPUs only touch synthetic and public data.
Base weights: frozen · Adapter: <1% of model · Training run: hours to days
Engagement Playbook
Every engagement runs the same way.
Standard steps mean predictable timelines, clean documentation for auditors, and no surprises.
Step 01
Data intake checklist
What documents exist, where they live, what format they’re in, and who owns access.
Step 02
PHI handling protocol
Signed documentation of exactly how data moves inside your environment during training. Nothing external.
Step 03
Dataset preparation
The hard part. We work with your team to turn messy real-world documents into clean training examples. Around 70% of the engagement.
Step 04
LoRA training on your hardware
Adapter training runs on your GPUs. Overnight or weekends. Hours to a couple of days of compute.
Step 05
Validation
The adapter runs against our evaluation suite. Improvement is measured. Regressions are caught before anything reaches production.
Step 06
Handover
Signed-off adapter, documentation for auditors, and a named person accountable for the artifact. The adapter belongs to you and stays inside your walls.