Input dataDocuments

Documents

The Data Room: a document repository where source files are stored, organised and read to extract environmental data.

Documents

The Documents module — the Data Room in the product — is the centralised repository for the source files behind an assessment: invoices, purchase records, supplier declarations, certificates, reports. It is where an input value's evidence lives, and where Darwin reads that evidence to propose data points.

What it holds

Files are uploaded directly or dragged onto the page, in the formats a company already has: PDF, Word, Excel, PowerPoint, CSV, Markdown and images. They are organised into folders — by topic, supplier, site, or whatever structure fits the workflow — and folders nest, so a large estate stays navigable. Two views read the same repository: Folders, to browse the hierarchy, and Search, to find a document across all folders by full-text search and filters.

The extraction pipeline

A document is not just stored. Once uploaded it moves through a sequence of states, visible on its row:

StateWhat is happening
UploadedThe file is stored and queued.
ProcessingIts content is being read and parsed.
EnrichingEnvironmental data is being extracted, and a summary and keywords derived.
ReadyExtraction is complete; the row expands to show the summary, the keywords and the extracted detail.

What extraction proposes reaches Data points as AI candidates — pending rows an analyst reviews and validates. Nothing extracted from a document enters an assessment without that validation step, which is what keeps the data quality score meaningful: the score reflects the type of input data, and an unreviewed extraction is not yet an input.

Documents are evidence, not an input type. They do not carry a footprint of their own. Their role is to justify a data point and to accelerate its entry — the assessment reads the data point, never the file.