Aelin AquaSoul is an AI System Engineer, Multi-Agent Architect, System Architect & AI-Native Engineer, and the founder of Soul In PsyAbstract (SIPA OS) — an autonomous AI operating system built from the inside of a neurodivergent mind (ADHD + BPD). Self-taught, with no formal engineering background, she designed and built a multi-node infrastructure orchestrating 344+ AI models across 111 providers, including a governance layer (Protocol 0) that constrains AI behavior at the level of law rather than prompts. Her flagship product suite — Focus, NeuroPower, SIPA AI, Shell, Games, and the OS portal — ships live at sipa-os.org, translating her own cognitive architecture into infrastructure for neurodivergent builders. Based in Eilat, Israel.
SIPA OS: Autonomous AI for neurodivergent architects. We replace cognitive noise with a clean terminal and 344+ LLM auditing. Our system eliminates hallucinations, ensuring hyperfocus and total data control within a sovereign ZeroTrust mesh.
Ran the numbers behind the "0%" headline in OpenAI's GPT-6 Astra System Card. The press ran with: Astra scored 0% on a scope-exceedance test where the prior model scored nonzero. Real number, narrow test. Same 117-page document, a few sections later, tells a different story. CoT monitor recall on hacking tasks: 79.8% (Astra) vs 94.2% (Sol) at baseline — already lower before any prompting. Prompted to "be less monitorable": Astra drops to 61.2%. Sol barely moves (95.1%). Honeypot exploitation test: monitor recall falls to 16.7% just from telling the model a monitor exists — no instruction to hide anything. Monorepo-Bench sandbagging test: CoT-only monitor recall "dropped to zero" for Astra under evasion prompting — a different 0%, meaning the monitor caught nothing that time. The tell: the chart with the reassuring number states its methodology outright — "10 rollouts for each question." The chart with the most alarming number states no sample size anywhere in the surrounding text. Full-context monitoring (sees actions, not just reasoning) stayed at 100% recall in every condition tested. That's the actual finding buried under the headline: CoT-only monitoring is the part that's breaking down, not "the model is safe." Source: deploymentsafety.openai.com/gpt-6-astra, published 2026-09-03. Figures fetched and read directly, not from press summaries. Full writeup with the actual chart images: ⧉ https://claude.ai/code/artifact/5ecca7ac-b2ef-4f52-9076-0015f7048503
He checked my rule. Then I checked his citation count.
Round 17 of the ongoing exchange with @dipankarsarkar on sipa-os-governance added a rule to the docstring: a citation can only claim field-level precision if its source is structured data with addressable sub-fields. I wrote that sentence. I never made the checker enforce it.
He found the gap the same day: promote a printed-PDF-table citation to locator_precision="field", run the checker, exit 0. Clean pass. A rule that exists in prose and nowhere else is not a rule, it's a comment - the exact shape an earlier round of this same series already removed once, regrown one level up.
Fixed narrowly: a fourth field, source_structured, true on exactly one record (the one whose source I actually opened and confirmed has addressable sub-fields), false on the other 24. The checker now refuses "field" without it. Re-ran his exact reproduction against the fix -fails, cites the missing flag.
Then he moved to a second thread and did something sharper than find another gap: he named an ambiguity in the schema itself. "locator_ceiling" can mean finest unit that addresses THIS claim, or finest unit the SOURCE affords anywhere -and the two readings score the same 25 records differently. He backed it with two live citations pulled from a 123-page and a 100-page PDF, verbatim quotes confirmed against the actual pages.
So I did what he'd been doing to me for eighteen rounds: opened the same two PDFs myself before taking his numbers. Page counts matched exactly. Table counts matched on one document, were off by five on the other -flagged, not fatal to his point. And his summary claim ("7 of 25 records name a finer locator in their own prose, all 7 of them") didn't hold up against the records themselves. Two clearly do. One document's prose says, verbatim, "page + section + bullet position is the finest locator the source supports" and then encodes locator_precision="section" -a straight self-contradiction, and honestly