When enterprises started putting AI assistants over their own content, a pattern showed up quickly. Some organizations had something useful running within weeks. Others spent a year on it and quietly stopped.
The difference was rarely the budget or the model. It was what state their documents were in when they started.
The ones who moved fast had done something unglamorous years earlier. They had put their documents into a managed repository, classified them on the way in, attached metadata, resolved which version of a thing was current, and set permissions that actually meant something. None of it was done with AI in mind. It just turned out to be the prerequisite.
A shared drive is not an archive
Most organizations believe they are already digital because their documents are files rather than paper. Then someone tries to build a retrieval system over those files and finds out what is actually there.
Eleven copies of the same policy, four of them edited after the version everyone treats as final. Filenames doing the work that metadata should be doing. Folder structures that made sense to a team that reorganized in 2019. Documents nobody can say who is allowed to read.
Retrieval over a corpus like that produces a specific failure that is worse than no answer. The assistant finds the superseded policy, quotes it accurately, and gives a confident response that is out of date. The user has no way to tell. This is the single most common reason these projects get shut down after a pilot.
Permissions are usually left until last
An assistant that answers from your documents inherits whatever access model the repository has. If the repository has none in practice, because everyone can reach everything on the shared drive, then the assistant will happily surface an HR investigation file or an unsigned board paper to whoever asks the right question.
Fixing that at the assistant layer is difficult and fragile. Fixing it at the repository layer is ordinary document management work that pays off whether or not AI is ever involved.
The expensive part is the cleanup
There is a common assumption that the cost of an AI project sits in the technology. In practice the model and the infrastructure are the predictable part. The unpredictable part is preparing the content, and it can consume most of the timeline.
Organizations that had already done that work spent their budget on the application. Organizations that had not found themselves running a document management project with an AI deadline attached to it, which is a harder thing to govern and a harder thing to explain to a sponsor who was promised an assistant.
Paper is still part of this
In this region a great deal of the material that matters most is still on paper, or exists as a scan that was never indexed. Contracts, personnel files, title documents, government submissions, decades of correspondence. Some of it is the only copy.
Scanning alone does not solve it. A folder of image files is no more searchable than the box it came from. The value appears when capture, classification and indexing happen together, so the document arrives in the repository already understood: what it is, who it belongs to, when it takes effect and how long it has to be kept.
That work has an obvious benefit before any AI is involved. It also happens to be the thing that decides whether AI can use the material later.
Starting now still counts as early
It is easy to read all of this as an argument that the moment has passed. It has not. Most enterprises are not in the ready group, and the ones already benefiting are a smaller club than the conference circuit suggests.
It also does not require doing everything at once. Start with the corpus people actually ask questions about. Policies, procedures, contracts, product documentation. Get that set into a managed repository with real metadata and real permissions, and you have something worth pointing an assistant at. The rest can follow on its own schedule.
The organizations that will get value from this over the next few years are the ones treating their archive as infrastructure rather than storage.
EDC handles both halves of that work: digitising physical records at scale, and putting the digital ones into a managed, governed repository that AI can actually use.

