How the work was built and measured
The engineering behind Nadhi Audit, and the on-device AI research it grew out of. If you want the product, it is on the Nadhi Audit page.
New: peer-reviewed in BMJ Digital Health & AIHigh-accuracy ECG image interpretation using parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning with multimodal LLaMA V.3.2
Nandakishor Mukkunnoth (Convai Innovations), Anjali Mukkunnoth (Dr. Moopens Medical College; AIIMS Rishikesh), Meenakshi Khapre (AIIMS Rishikesh), Sudarshan Benamanahalli Puttaswamy (Dr. Moopens Medical College)
We fine-tuned an open multimodal model with LoRA on ECGInstruct, one million ECG images paired with expert annotations, so it reads an ECG image and writes the report. The aim is clinical ECG support where a cardiologist is not at hand.
The same recipe is behind Nadhi Audit: take an open model, fine-tune it cheaply for one specialist job, and run it where the sensitive data already is.
Nadhi Audit: how it works and how it was measured
Nadhi Audit is an audit harness driven by a model we fine-tuned ourselves, Nadhi_Audit_FT.gguf, quantised to run on an Apple Silicon Metal GPU. Deterministic checks run first and hand the model a shortlist; the model reads each candidate in context, maps it to a CWE, and writes the patch. Nothing is sent anywhere, including the advisory lookups: the whole CVE database is downloaded and every comparison happens locally, so we never learn which packages a customer uses. The database is published in the open and rebuilt daily.
How the model was trained
The base is Gemma 4 E4B, fine-tuned with bf16 LoRA (r=32, alpha=32) at a 12,288 token context. Two things had to be learned together, because an auditor that can recognise a vulnerability but cannot drive its own tools is no use: the security work itself, and agentic tool calling. Roughly half the set is audit and compliance work, and half is generic multi-step tool use.
Every conversation was generated in batch and then ranked by a judge model, and the bottom of the ranking was thrown away rather than trained on. The thirty behaviour tracks are deliberately not all happy paths. The largest single track is a sequential chain of three to six tool calls, and the set also carries parallel dispatch, state carried across turns, pronoun resolution across a follow-up, recovery after a tool returns an error, refusing to call anything when a required parameter is missing, stopping instead of retrying a call that keeps failing, and answering from what is already known without calling a tool at all. About 13,000 of the conversations are adversarial. Every track is represented on macOS, Linux and Windows in roughly equal measure, and the tool call count per conversation runs from zero to ten.
The most instructive bug in the whole training run was the loss mask. Stock response-only masking, applied to Gemma 4's chat template, leaves the tool output block inside the trained span, which teaches the model to write the result of a tool call itself instead of stopping and waiting for one. A model that invents its own tool output is worse than useless in an auditor, because it will invent findings. A custom mask over the template fixed it.
Tool selection came out reliable. Argument grammar was the weaker half, which is why the harness validates every argument against the tool schema before dispatch rather than trusting what the model emitted. That is the same principle as the fix verifier below: the model proposes, and something deterministic decides.
The checks
Injection, cross-site scripting, weak crypto and secrets, plus the mistakes specific to generated code: a service-role key behind a public environment variable, authorisation decided in the browser, a model API key shipped to the client.
Replays the database migrations in order and reports which tables end up served with no access policy at all.
Every version the lockfile resolved, for npm, pnpm, yarn, bun, uv, poetry and composer, matched offline against a bundled OSV export. Names that never resolved are reported too, because that is what a hallucinated package looks like.
Optional and read-only. The migrations say what was intended; this checks what the live project actually enforces.
Sixteen checks run over one tree and their results are correlated, so findings that combine across lanes are reported as one chain rather than four separate lines. Every report names which checks ran, what each raised, and what was discarded before review.
Why the fixes are verified twice
The judgement the model cannot make is whether its own patch worked. Across four repositories, sixteen proposed fixes all applied cleanly, only about seventy percent actually removed the finding, one left a file that would not parse, and the pass reported every one of them as fixed. The verifier exists because of that measurement. Each patch is now re-parsed by the real compiler (node --check, py_compile, php -l), the rule that raised the finding is re-run against the patched file, and a patch that fails either check is held rather than reported as done. Code edits are off by default; when they are on, clean fixes land on an isolated nadhi branch so they stay a diff you review.
The 30-repository run
Thirty public repositories nobody prepared for it, 71,418 source files, 836 findings. Every dependency finding was checked against osv.dev, the upstream advisory database, rather than the offline copy the product ships with, so the check is not circular. 799 of 802 were confirmed, which is 99.6 percent.
| Check | Found |
|---|---|
| Known vulnerable dependencies | 805 |
| Weak cryptography | 4 |
| Error disclosure | 3 |
| Path traversal | 2 |
| Authentication | 2 |
| Hardcoded secrets | 2 |
| Personal data exposure | 2 |
| Script integrity | 2 |
| Database privileges | 2 |
| Cross-site scripting | 1 |
| Transport security | 1 |
| Row-level security | 0 |
| Public env exposure | 0 |
| Injection | 0 |
| SSRF | 0 |
| Deserialization | 0 |
| Total | 836 |
Cost of the same work in the cloud
Token cost of auditing the same code once a week for a year, priced at each vendor's published rate against the token count this run actually consumed. Nadhi Audit consumed zero API tokens, because the model runs on the machine doing the audit.
| Engine | Per year |
|---|---|
| GPT-6 Astra | $72,807 |
| Claude Opus 5 | $36,404 |
| GPT-5.6 Sol | $29,123 |
| Gemini 3.8 Flash | $5,461 |
| Nadhi Audit, 5 machines, unlimited repositories | $600 |
What we have not measured
- Recall. On a real repository you cannot know what was missed, so the honest position is that we report precision and do not claim a recall figure.
- Lockfiles for npm, PyPI and Packagist are read. NuGet, Go, Maven and RubyGems are not yet, and a project built on those gets no dependency findings at all.
- There is no SBOM output yet.
Statutory citations in the DPDP, HIPAA and GDPR report modes point at the section each finding engages. They are a technical mapping, not a legal opinion.
AI4Cardio: on-device multimodal cardiac AI
AI4Cardio is an on-premise diagnostic system built with cardiology specialists. Multimodal vision-language models are fine-tuned with LoRA on 12-lead ECG waveforms and patient clinical parameters, giving arrhythmia detection and cardiac risk stratification on a clinic workstation.
Inference runs entirely on the local GPU, so patient records and raw physiological telemetry never leave the hospital, which is what the residency requirements in HIPAA and DPDP actually ask for.
Nadhi: an autonomous desktop co-scientist
Nadhi Co-Scientist is an on-device research assistant. Specialised sub-agents (planner, critic, director, synthesiser) form hypotheses, fetch and index hundreds of open-access papers from arXiv, PMC, PubMed and bioRxiv, run local Python experiments, and write cited manuscript drafts in DOCX and PDF with full revision history.
Our confidence-routing work is what keeps it grounded: every generated paragraph has to trace back to a verified citation before it is written.
Papers and preprints
Confidence-aware routing for hallucination mitigation
Multi-signal confidence routing that suppresses a hallucination before generation rather than filtering it afterwards.
Read on arXiv →High-accuracy ECG image interpretation with LoRA fine-tuning of multimodal LLaMA V.3.2
AUC 0.98 against 0.51 for the base model, trained on one million annotated ECG images.
DeepRAG: custom embedding models from scratch
Embedding architectures for local document retrieval with no third-party vector cloud in the path.
Read on arXiv →Continual learning agents via A2C reinforcement learning
Advantage actor-critic training for agents that keep adapting on the device they run on.
Read on arXiv →