Transparency in public procurement is usually discussed as a question of whether a document was published. There is a second question that is almost never asked: published in what form.
Public, but unreadable
A large share of decisions and evaluation protocols are not uploaded as text. They are uploaded as a picture of text — a scanned page, sometimes signed and stamped, saved as an image inside a PDF.
For someone who has already found the document, the difference is nil. They open it and read.
For anyone searching, the difference is absolute. A scan has no text layer. It is not found by keyword, not indexed, never appears in a query result. If the exclusion ground you care about is described only in a document like that, it effectively does not exist as far as search is concerned.
The result is a quiet split: documents that are public and findable, and documents that are public and invisible. The second group is not small.
What surfaces once they are read
Reading is not uniform. The newer generation of processing extracts not only the text but the structure — tables stay tables, headings stay headings. That matters, because much of the useful content in an evaluation protocol sits precisely in a table: which bidder, on which criterion, what score, on what ground.
Flat text is good enough for finding. Structured text is what analysis needs. So 367,557 of the read documents went through the newer pass, while the older 293,004 remain flat text until they are reprocessed.
Why this is an advantage, not a technical footnote
When someone says “we analysed exclusion grounds from public data”, the first fair question is: including the scanned decisions, or only the machine-readable ones?
If only the machine-readable ones, the analysis rests on a systematically skewed sample. Not random — skewed. Buyers who publish scans do so consistently rather than occasionally. Which means entire buyers and entire procedure types fall out of the picture.
That is the difference between “we have no data on this buyer” and “this buyer uploads scans”. The first is a gap. The second is a solvable problem.
What reading does not fix
It does not fix the quality of the document itself. If a protocol is vaguely worded, reading it produces vague text that can at least be found.
Nor does it fix absence. Around 7% of tenders have no attached document at all — there is nothing there to read.
But for everything else, a scan stops being the end of the trail. And in work whose whole point is to find the precedent before you repeat someone else’s mistake, that is the difference between an answer and a guess.