Research and Engineering Notes

How the work was built and measured

The engineering behind Nadhi Audit, and the on-device AI research it grew out of. If you want the product, it is on the Nadhi Audit page.

New: peer-reviewed in BMJ Digital Health & AI
Peer-reviewedBMJ Digital Health & AI · Volume 1, Issue 1 · December 2025

High-accuracy ECG image interpretation using parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning with multimodal LLaMA V.3.2

Nandakishor Mukkunnoth (Convai Innovations), Anjali Mukkunnoth (Dr. Moopens Medical College; AIIMS Rishikesh), Meenakshi Khapre (AIIMS Rishikesh), Sudarshan Benamanahalli Puttaswamy (Dr. Moopens Medical College)

We fine-tuned an open multimodal model with LoRA on ECGInstruct, one million ECG images paired with expert annotations, so it reads an ECG image and writes the report. The aim is clinical ECG support where a cardiologist is not at hand.

0.98
AUC
base model 0.51
0.74
Macro F1
base model 0.33
85.4
Report quality score
base model 47.8
>90%
Common arrhythmias
accuracy

The same recipe is behind Nadhi Audit: take an open model, fine-tune it cheaply for one specialist job, and run it where the sensitive data already is.

🛡️
Product Engineering

Nadhi Audit: how it works and how it was measured

Nadhi Audit is an audit harness driven by a model we fine-tuned ourselves, Nadhi_Audit_FT.gguf, quantised to run on an Apple Silicon Metal GPU. Deterministic checks run first and hand the model a shortlist; the model reads each candidate in context, maps it to a CWE, and writes the patch. Nothing is sent anywhere, including the advisory lookups: the whole CVE database is downloaded and every comparison happens locally, so we never learn which packages a customer uses. The database is published in the open and rebuilt daily.

How the model was trained

The base is Gemma 4 E4B, fine-tuned with bf16 LoRA (r=32, alpha=32) at a 12,288 token context. Two things had to be learned together, because an auditor that can recognise a vulnerability but cannot drive its own tools is no use: the security work itself, and agentic tool calling. Roughly half the set is audit and compliance work, and half is generic multi-step tool use.

71,152
trajectories generated
40,000
kept after ranking
1,426
held back for validation
30
behaviour tracks

Every conversation was generated in batch and then ranked by a judge model, and the bottom of the ranking was thrown away rather than trained on. The thirty behaviour tracks are deliberately not all happy paths. The largest single track is a sequential chain of three to six tool calls, and the set also carries parallel dispatch, state carried across turns, pronoun resolution across a follow-up, recovery after a tool returns an error, refusing to call anything when a required parameter is missing, stopping instead of retrying a call that keeps failing, and answering from what is already known without calling a tool at all. About 13,000 of the conversations are adversarial. Every track is represented on macOS, Linux and Windows in roughly equal measure, and the tool call count per conversation runs from zero to ten.

The most instructive bug in the whole training run was the loss mask. Stock response-only masking, applied to Gemma 4's chat template, leaves the tool output block inside the trained span, which teaches the model to write the result of a tool call itself instead of stopping and waiting for one. A model that invents its own tool output is worse than useless in an auditor, because it will invent findings. A custom mask over the template fixed it.

Tool selection came out reliable. Argument grammar was the weaker half, which is why the harness validates every argument against the tool schema before dispatch rather than trusting what the model emitted. That is the same principle as the fix verifier below: the model proposes, and something deterministic decides.

The checks

01
Source patterns

Injection, cross-site scripting, weak crypto and secrets, plus the mistakes specific to generated code: a service-role key behind a public environment variable, authorisation decided in the browser, a model API key shipped to the client.

02
Row-level security

Replays the database migrations in order and reports which tables end up served with no access policy at all.

03
Supply chain

Every version the lockfile resolved, for npm, pnpm, yarn, bun, uv, poetry and composer, matched offline against a bundled OSV export. Names that never resolved are reported too, because that is what a hallucinated package looks like.

04
Deployed backend

Optional and read-only. The migrations say what was intended; this checks what the live project actually enforces.

Sixteen checks run over one tree and their results are correlated, so findings that combine across lanes are reported as one chain rather than four separate lines. Every report names which checks ran, what each raised, and what was discarded before review.

Why the fixes are verified twice

The judgement the model cannot make is whether its own patch worked. Across four repositories, sixteen proposed fixes all applied cleanly, only about seventy percent actually removed the finding, one left a file that would not parse, and the pass reported every one of them as fixed. The verifier exists because of that measurement. Each patch is now re-parsed by the real compiler (node --check, py_compile, php -l), the rule that raised the finding is re-run against the patched file, and a patch that fails either check is held rather than reported as done. Code edits are off by default; when they are on, clean fixes land on an isolated nadhi branch so they stay a diff you review.

The 30-repository run

Thirty public repositories nobody prepared for it, 71,418 source files, 836 findings. Every dependency finding was checked against osv.dev, the upstream advisory database, rather than the offline copy the product ships with, so the check is not circular. 799 of 802 were confirmed, which is 99.6 percent.

CheckFound
Known vulnerable dependencies805
Weak cryptography4
Error disclosure3
Path traversal2
Authentication2
Hardcoded secrets2
Personal data exposure2
Script integrity2
Database privileges2
Cross-site scripting1
Transport security1
Row-level security0
Public env exposure0
Injection0
SSRF0
Deserialization0
Total836

Cost of the same work in the cloud

Token cost of auditing the same code once a week for a year, priced at each vendor's published rate against the token count this run actually consumed. Nadhi Audit consumed zero API tokens, because the model runs on the machine doing the audit.

EnginePer year
GPT-6 Astra$72,807
Claude Opus 5$36,404
GPT-5.6 Sol$29,123
Gemini 3.8 Flash$5,461
Nadhi Audit, 5 machines, unlimited repositories$600

What we have not measured

  • Recall. On a real repository you cannot know what was missed, so the honest position is that we report precision and do not claim a recall figure.
  • Lockfiles for npm, PyPI and Packagist are read. NuGet, Go, Maven and RubyGems are not yet, and a project built on those gets no dependency findings at all.
  • There is no SBOM output yet.

Statutory citations in the DPDP, HIPAA and GDPR report modes point at the section each finding engages. They are a technical mapping, not a legal opinion.

🫀
Clinical Research

AI4Cardio: on-device multimodal cardiac AI

AI4Cardio is an on-premise diagnostic system built with cardiology specialists. Multimodal vision-language models are fine-tuned with LoRA on 12-lead ECG waveforms and patient clinical parameters, giving arrhythmia detection and cardiac risk stratification on a clinic workstation.

Inference runs entirely on the local GPU, so patient records and raw physiological telemetry never leave the hospital, which is what the residency requirements in HIPAA and DPDP actually ask for.

Architecture
Multimodal LoRA
12-lead ECG waveform analysis
Privacy
Fully on-premise
No patient data leaves the site
Peer-reviewed
BMJ Digital Health & AI
December 2025 · AUC 0.98
🔬
Autonomous Science

Nadhi: an autonomous desktop co-scientist

Nadhi Co-Scientist is an on-device research assistant. Specialised sub-agents (planner, critic, director, synthesiser) form hypotheses, fetch and index hundreds of open-access papers from arXiv, PMC, PubMed and bioRxiv, run local Python experiments, and write cited manuscript drafts in DOCX and PDF with full revision history.

Our confidence-routing work is what keeps it grounded: every generated paragraph has to trace back to a verified citation before it is written.

Literature
100+ open papers
arXiv, PMC, OpenAlex
Output
DOCX and PDF
Format-preserved edits with undo
Method
Multi-agent planner
Empirical code validation loop

Papers and preprints

arXiv:2510.01237

Confidence-aware routing for hallucination mitigation

Multi-signal confidence routing that suppresses a hallucination before generation rather than filtering it afterwards.

Read on arXiv →
BMJ Digital Health & AI · 2025 · peer-reviewed

High-accuracy ECG image interpretation with LoRA fine-tuning of multimodal LLaMA V.3.2

AUC 0.98 against 0.51 for the base model, trained on one million annotated ECG images.

arXiv:2503.08213

DeepRAG: custom embedding models from scratch

Embedding architectures for local document retrieval with no third-party vector cloud in the path.

Read on arXiv →
arXiv:2502.12876

Continual learning agents via A2C reinforcement learning

Advantage actor-critic training for agents that keep adapting on the device they run on.

Read on arXiv →