Machines generate. Judgment is Human.
When producing something becomes nearly free, what does a human still have to do?
01 / the lab
An independent lab for the judgment problem.
Mudhawi Labs is an independent AI lab. We work on one problem: what it takes to trust what artificial intelligence produces. The tools that generate software, content, and decisions have outrun the tools that judge them, and nearly everyone now owns something they cannot inspect.
The lab is independent by construction. It sells no software, holds no position in the systems it examines, and takes no retainer from anyone whose work it might one day assess. Independence is not a claim here. It is the business model.
The lab also runs on its own thesis. AI systems handle our research support, our operations, and our correspondence. They never write a finding. Generation is delegated here too, because that is the sane default now. The reading and the judgment are not, and every finding carries the name of the person who reached it.
Research is the lab’s centre of gravity. Where the research meets running systems, it becomes a practice, and today that practice is the audit.
02 / what the lab studies
One question. Three programmes.
Producing things has become nearly free. Writing software, drafting documents, generating analysis: the cost of making has collapsed, and it is not coming back. What has not become cheaper is knowing whether what came back is any good.
That gap is the lab’s subject. When generation is free, what remains a human obligation, and how can that obligation be discharged in a form somebody else can check? The programmes below are that question pointed at three different targets.
Verifying AI-built systems
The question aimed at code. A founder with no engineering background can describe a product in plain language and have a working version by the end of the week. What has not become cheaper is the judgment required to know whether the thing that came back does what it claims, holds when a stranger pushes on it, and can be maintained by whoever inherits it.
Three problems sit unresolved in that gap. Provenance: when a system is assembled largely from generated code, nobody can say with confidence which decisions were deliberate and which were incidental. Coverage: generated tests tend to confirm the behaviour the generator already assumed, so the tests pass and the assumption is never examined. Drift: a system that was correct on the day it was built stops being correct quietly, as its dependencies, its models, and its data move underneath it.
AI and entrepreneurship
The question aimed at company-building. When production costs collapse, the constraint on starting something moves. It stops being the ability to build and becomes the ability to judge: which of the things you can now make cheaply is worth making, and how would you know before the market tells you.
The lab studies what founders are actually deciding under these conditions, what they delegate to machines, what they discover they cannot, and what a defensible decision looks like when the cost of producing an option has fallen to nearly nothing. This is the founder’s doctoral subject and the lab’s longest-running line of work.
Arabic AI, evaluated
The question aimed at a language. Arabic-language systems are being deployed faster than the field can measure them. Most public benchmarks were designed for English and then translated, which tests whether a model can handle a translation rather than whether it can handle Arabic as it is written and spoken. The material that survives translation is the material that was easiest to translate, so scores are highest exactly where the language is least itself.
Dialect makes this worse, and it makes it interesting. A model answering in Modern Standard Arabic to a question asked in Khaleeji has produced a technically correct response and a socially wrong one, and there is no widely agreed way to score that. Diacritics carry meaning the written form usually omits, so two readings of the same string can both be defensible.
This matters because measurement decides procurement. A benchmark that is quietly unsuitable does more damage than no benchmark at all, because it converts an open question into a decided one.
The lab does not build Arabic language systems and has no interest in doing so. It studies how they are evaluated, benchmarked, and trusted.
03 / lab notes
We publish what we work out.
The first notes are being written.
Lab notes is where the lab thinks in public: what the research is turning up, what the practice keeps running into, and which widely repeated claims about AI do not survive being checked. Written to be read by people who are not specialists, and to be checkable by people who are.
Nothing is listed here yet because nothing has published yet. A list of titles with no writing behind them is exactly the kind of claim this lab exists to catch.
04 / follow the lab
The work is published as it happens.
Teardowns, notes from the practice, and what the research turns up. Same standard as everything else here: claims carry receipts, and corrections are read and credited.
Nothing is posted to a channel the lab would not sign.
05 / the practice
Where the research meets running systems.
The audit
An independent review of a product built with AI, scoped and priced in full before it begins, delivered as a signed report. It is the applied face of the first research programme, and the only commercial service the lab offers.