For hospital IT, compliance & revenue-cycle teams
On-Premise AI That Keeps Patient Data Inside Your Hospital
We install open-weight models on your hardware, tune them for one workflow at a time, and log every output for human review. No external model API.
Now accepting pilot hospitals · core platform and PDF processing available as working prototypes
Llama · Qwen · Gemma · Phi · Mistral · vLLM · LoRA · MCP
What is GofarAI Labs?
GofarAI Labs designs and deploys open-weight AI systems for US hospital workflows. Models run on hospital hardware, with patient data, inference, and audit trails kept inside the hospital’s environment.
Who it’s for
Built for the teams buried in documents.
Revenue cycle
Draft appeal and justification letters from chart notes. Your team edits, approves, and sends.
Denial analysis · Prior auth · Coding assistance
Health information management
Sort inbound faxes and referrals, summarize long records, and process large PDF stacks page by page.
Fax triage · Chart summaries · PDF batches
Compliance & quality
Every inference is logged with who asked, when, and what data was in scope — ready for an audit.
Audit logs · Audit preparation · 42 CFR Part 2
IT & security
Runs on your hardware behind your egress firewall, with no required outbound calls and one gateway for AI agents.
On-prem GPUs · RBAC · MCP gateway
The Thesis
Hospitals need control over how AI handles PHI.
Using cloud LLMs with PHI can require approved vendors, business associate agreements, security review, data-governance controls, and ongoing oversight. For hospitals that need tighter operational control, GofarAI offers an on-premise alternative.
We deploy open-weight models inside the hospital environment and tune them for defined workflows. This keeps inference under hospital control, removes the need for an external model API during normal operation, and makes model activity auditable.
How data flows
Every step happens inside your walls.
- 01Documents & requestsFaxes, PDFs, chart notes, EHR exports
- 02GofarAI serverOpen-weight model on your GPU. Every call audit-logged.
- 03Draft outputSummary, routing, letter, or code suggestions
- 04Human reviewYour staff approve, edit, or reject
The Stack
The full stack, installed on your premises.
Five layers, one deployment. Every layer designed to keep patient data inside your building.
Select a layer to see what’s inside
Layer 05MCP GatewayOne inspectable control point for every AI agent in your hospital.MCP serverMCP clientPolicy
- Exposes GofarAI tools as MCP endpoints, so authorized agents — EHR-adjacent systems, other vendors’ agents, future AI purchases — can call them through a standard protocol.
- Acts as an MCP client, letting the local model call other hospital systems that expose MCP endpoints, such as scheduling, labs, or bed management.
- Enforces policy per agent: which tool it can call, with what data scope. Every call is logged.
What you control: You decide which agents get in and what data they can see.
Layer 04ToolsModular workflow plugins. Start with one, add more over time.PluginsHuman approval
- Every tool follows the same shape: read → extract → classify → draft.
- Same server, same models — a new tool means new prompts and output templates, not new infrastructure.
- Start with document intake; add revenue-cycle and clinical documentation tools as you go.
What you control: Your staff approve every output before it is used.
Layer 03ModelsTuned open-weight models in the 7B–14B range, swappable as better ones ship.7B–14BLoRA adaptersRouter
- Open-weight families such as Llama, Qwen, Gemma, Phi, and Mistral — never locked to one vendor’s roadmap.
- Larger hospitals can run role-specialized models: one for revenue-cycle language, another for clinical documentation.
- Department isolation keeps fine-tunes separate so weights never mix information domains — important for 42 CFR Part 2 data. A small router model sends each request to the right one.
What you control: Adapters trained on your data belong to you and stay on your hardware.
Layer 02PlatformThe installable GofarAI package — where the HIPAA-friendly architecture lives.vLLMRBACAudit logEncryption
- Inference server (vLLM or llama.cpp) and a model manager for swapping models.
- Audit logging of every inference, prompt, and completion.
- Role-based access controls and encryption at rest.
- Zero external calls — the platform never phones home.
What you control: It runs behind a strict egress firewall; updates arrive as signed offline packages on your schedule.
Layer 01HardwareWe spec, source, and install the server and GPU — or certify hardware you already own.24GB GPUOn-prem server
- A single 24GB GPU runs a quantized 8–14B model well for document workflows.
- Larger deployments get multi-GPU servers with department isolation between models.
- Speed is not the selling point. Privacy and local control are.
What you control: The server sits on your network, under your control.
- 1 GPU
- A single 24GB card runs an 8–14B model for document work
- 0
- Required outbound calls during normal operation
- Every
- Inference, prompt, and completion is audit-logged
Tool Modules
Workflows your teams use every day.
Read → extract → classify → draft. Same engine, different prompts. A human approves every output.
Working prototypes are ready for pilot evaluation. Pilot concepts are co-designed with partner hospitals.
Working prototype
PDF Batch Processing
Page-by-page OCR, structure extraction, and summarization across large document stacks.
Pilot concept
Fax & Referral Triage
Classify inbound documents, extract key fields, route to the right department queue.
Pilot concept
Prior Authorization Drafting
Assemble clinical justification letters from chart notes. Human clinician reviews and signs.
Pilot concept
Denial Letter Analysis
Extract denial reason codes, suggest appeal grounds, draft appeals for the revenue cycle team.
Pilot concept
Coding Assistance
Suggest ICD-10 and CPT codes from encounter notes for human coders to verify.
Pilot concept
Discharge Summary Drafting
Assemble discharge summaries from chart data. Physician reviews and signs off before release.
Pilot concept
Chart Summarization
Condense transferred records before appointments so clinicians walk in prepared.
Pilot concept
Audit Preparation
Search and compile documentation for Joint Commission and payer audits in hours, not weeks.
Open Source Position
Open where it should be. Closed where it must be.
A clean legal and philosophical line, drawn in the same place every time.
What we plan to open-source
- Training recipes and tuning pipelines
- Evaluation benchmarks for healthcare tasks
- Synthetic and de-identified datasets
- Deployment tooling and platform code
- Base model weights tuned on public data
We never open-source
- Anything trained on real patient data
- Hospital-specific model adapters
- Hospital-specific evaluation results
- Anything that could reidentify individuals
- Weights that ever touched PHI
Security & Auditability
Every inference. Every interaction. Inspectable.
No outbound service dependency during normal operation. Prompts, completions, and tool invocations are designed to be logged inside your building. When other vendors’ AI agents show up in your hospital, the MCP gateway provides a control point for which agent can call which tool, with what data scope, all reviewable.
Fine-Tuning as a Service
We tune. On your hardware. On your data.
On narrowly defined workflows, a smaller model tuned and evaluated against local requirements can outperform a larger general-purpose model. We train LoRA and QLoRA adapters on your existing GPUs, overnight or on weekends. Training is designed to occur inside the hospital environment without transferring training data to an external model provider.
Build the AI substrate for your hospital.
A short call. We map your workflows to a deployment plan and walk through the security model and the pilot plan.