Exhibit DAI security
AI proposes, evidence decides
- 7
- AI-assisted techniques, each tied to the evidence that must exist before its output counts
- 3
- defenses on one federated model: poisoning filter, differential privacy, drift correction
- 5
- stages before an automated fix runs: simulate, verify, approve, execute, health-check
D.1 Threat model: letting a language model near systems I run
In plain English: a map of how an AI assistant connected to my own systems could be misused, and the safeguard I put in place at each step.
Two of my own systems use a language model: my operations platform, HoWz, has a question pane where a model may suggest changes, and the lab runs a small local model on a CPU. Pick any part of the flow to see what can go wrong there and the control in the design.
Illustrative Written for this page from the design notes of HoWz and the lab. It is not an audit.
Threats at this step
Hover, tap or tab to a step in the flow. Each shows the threats that enter there and the control in the design, with its source.
The threat model as a table
| No. | Threat | Enters through | Control in the design | Source |
|---|---|---|---|---|
| T-1 | Prompt injection steers the model toward a harmful change | Text the model reads | The model can only propose. Code validates every proposal, and nothing is written until I approve it. | HoWz |
| T-2 | Excessive agency: the model acts instead of suggesting | Model output | Proposals are data, not actions: nothing is written until a separate human step approves it. | HoWz |
| T-3 | Tampered model runtime | Installation | Installed only after its release checksum was verified. | Lab |
| T-4 | Model endpoint reachable from outside | Network | The local model listens on the lab network only. | Lab |
| T-5 | Stolen session on the approval screen | Browser | Password gate, Content Security Policy, and HttpOnly, SameSite=Strict cookies. | HoWz |
D.2 The rule, in the research
In plain English: in my research an AI tool may point at a possible problem, but the problem only counts once a real test confirms it.
Language models are fast and confidently wrong often enough that I do not let one close a finding. In the Breakwater paper every AI-assisted technique gets the same treatment: it may propose candidates and rank them, but a finding closes only when a measured or scoped, simulated witness exists. AutoML ranks known CVE features; it does not discover vulnerabilities. An LLM proposes fuzzing inputs; the proof is the crash, not the model’s description of one.
| Technique | What it may do | Witness required |
|---|---|---|
| AutoML triage ranking | rank candidates | measured CVE and EPSS features |
| Nuclei templates | targeted checks | measured template match |
| OpenVAS / GVM | scanner evidence | measured plugin result |
| Default-credential check | scoped proof | simulated loopback login |
| Firmware entropy | supply-chain signal | measured file evidence |
| Protocol grammar from captures | structure inference | measured wire capture |
| LLM mutation and RL fuzzing | propose test inputs | measured crash or response |
D.3 AI as the target: defending the models themselves
In plain English: AI models can be attacked too, for example by feeding them bad training data. These are defenses I built and tested in my doctoral project.
A federated intrusion detector that expects some of its clients to lie. In Breakwater’s collaborative phase I trained a Transformer-based intrusion-detection model across simulated sites without pooling their raw traffic. Multi-Krum aggregation drops poisoned updates from malicious clients, differential privacy (Gaussian mechanism with Rényi accounting) bounds what any one site’s data can leak, and SCAFFOLD corrects client drift.
Autonomous agents on a short leash. A reinforcement-learning agent (PPO) plans offensive tests inside a decision model, but only behind a tiered safety controller that moves from simulation or shadow mode to controlled to autonomous, with every step written to a SHA-256 evidence chain. The remediation phase asks before it acts: simulate, verify, approve, execute, health-check, with rollback checkpoints.
Also on file: the App Academy certificate in AI-powered software development and generative AI engineering, Introduction to Generative AI (Google Cloud), the SANS AI Cybersecurity Forum, the Agentic AI Bootcamp at the Howard AI Network, a 2026 preprint on adversarial machine learning against automotive attack surfaces, and AI security risk and governance advisory at Cyntraix.