Picture a lab with twenty years of output on a shared drive. Published papers, bench protocols, a couple of theses, and the standard operating procedures nobody can find at the moment they need them.

Every new postdoc asks the same six questions in their first month. The answers exist. They are spread across eleven PDFs that nobody has read end to end since 2019.

This is a good fit for Docutrain, because the material is written down and the questions people ask about it are answerable from the text. What follows is the order a group would actually do it in.

Getting the PDFs in

The first decision arrives the moment you select a file: What type of PDF is this?

For a journal article or a protocol, the answer is Document, which pulls individual figures, tables and charts out of the pages so they can appear inline in answers. Slide deck captures whole pages instead, which is right for a presentation and wrong for a paper.

The second thing worth knowing is what a scanned PDF costs you.

A digital PDF has real text underneath and extraction is quick. A scan is a picture of a page, so Docutrain has to read it visually, which is the slowest and most fragile step in the whole pipeline. It is handled rather than refused: if the service doing that reading stalls or returns nothing, the same file goes to an alternative, with a plain text reader as a final fallback.

So a protocol photocopied in 1998 will train. It just takes longer, and you should not be alarmed by the wait. If a paper contains flowcharts or decision trees, an extra pass runs to read them properly, and the processing card tells you that is what it is doing.

Page count is the one hard limit: a warning past 300 pages, a block at 500. The fix is in the message. Split the file, train on the first part, then add each remaining part through Retrain with Add to existing data. The document keeps its link and every setting, and only the content grows.

Four settings that make a paper behave like a paper

The abstract. Generate one from the document's processed content, then switch on the Abstract tool so readers get a chip that opens it, complete with the date it was produced and a PDF export. If you retrain later, it flags itself as possibly stale rather than quietly serving a summary of a version that no longer exists.

References, which are not citations. This distinction matters more in research than anywhere else. Citations are generated automatically from your document's text and appear under answers as chips. References is a bibliography you write by hand, one entry per line, with URLs and DOIs turned into links automatically.

While you are there, check the Citations toggle is on. With it on, selecting a chip in an answer opens the verbatim source passage with the page it came from, and in multi-document chats the document too. For a research audience that viewer is the entire basis of trust: it is the difference between an assistant that claims your protocol says something and one that shows you the paragraph.

The PubMed link. For a document based on a published article, recording its PubMed identifier adds a button to the chat that fetches the article's title, authors, journal, year and abstract live, with a link through to the full record. It costs one field and saves every reader a search.

Figures. Everything extracted during processing lands on the Images tab with an AI-written caption and a longer description. The caption is what readers see under a figure; the description is what the assistant searches. Both are editable, and editing refreshes the index.

One control is worth knowing about specifically for research: Replace swaps in a better scan of a chart while keeping the caption, description and index intact. A clearer version of figure 3 does not cost you the work you did describing it. See images and video.

Grouping the body of work

One paper is one assistant. A body of work wants a collection.

Tick what belongs, then drag the documents into sequence. That order is exactly what readers see in the sidebar, so a deliberate arrangement, background papers then methods then protocols, is worth the two minutes it takes.

Allow chat across all documents is what turns a list of papers into a corpus. With it on, a reader can ask one question of everything, or tick a subset. And citations in a multi-document answer name which document each source came from, so "where does that number come from" stays answerable even when the answer drew on four papers at once.

Access is set on the collection itself, and one collection passcode unlocks every document inside it, including any whose own level is stricter. The editor warns you before you save. For a lab handing a reading pack to a collaborator, that is the behavior you want.

Reading what people actually asked

The part most groups underestimate is the Intelligence tab. It unlocks at five questions.

Keyword density counts the literal words and phrases people type, and it works from the very first question. Frequently the most useful thing it tells you is that your readers use different vocabulary from your papers.

Once your organization has around 200 questions, the Content gaps list ranks topics by a gap score built from answers that found no source, got re-asked, or were thumbed down, weighted by how many people asked. For a lab, that list is usually a straightforward instruction: a recurring question your protocols do not answer means a protocol needs a section written.

Two habits make this useful rather than decorative.

Turn on Exclude owner & admin traffic once the document is properly shared. Before that, most of the data is your own group testing it, and reading your own probing as audience behavior will send you somewhere unhelpful.

And use Doc, the analyst in the Intelligence panel. It is grounded in the question list rather than in your documents, and its suggested prompt "Where are the content gaps?" is the right first question. It quotes real examples and separates what the data shows from what it is inferring, which matters to an audience that will notice the difference. See document intelligence.

If you want to understand why the answers stay tied to your text in the first place, how a grounded answer is built is the place to start.