Half of what a document actually says often lives in its figures. The clinical algorithm. The org chart. The graph showing the trend the surrounding paragraph only gestures at.

An assistant that can only quote prose gives you a description of the flowchart. Docutrain extracts the figures during training, has a vision model describe each one, and puts the relevant image into the answer alongside the text.

Getting the pictures out

What comes out depends on the file, and mostly on one choice you make at upload.

SourceWhat is extracted
PDF as DocumentIndividual figures, tables, charts and photos. Graphics under about 5 KB are skipped as decoration, duplicates stored once, figures re-rendered at higher resolution where that is sharper.
PDF as Slide deckEvery page as a full slide image, trimmed to the slide's own edges where the export added letterboxing.
PowerPoint (.pptx)Text only. For the slides themselves, export to PDF and upload as a slide deck.
Admin uploadsAdded straight to the image library from the editor.

That Document versus Slide deck question is the most consequential image decision you will make, and it is asked before anything is processed. Reports and handbooks want Document. Presentation decks want Slide deck. The full flow is in training your first assistant.

Slide extraction runs in the background after the text is done, so you may see a note that the gallery is still being built while the assistant is already answering questions. Nothing is blocked on the pictures.

Every extracted image is then examined by a vision model, which writes a short caption and a longer description. The distinction matters and is worth memorizing: the caption is what readers see, the description is what the assistant reads.

Editing those two fields does more for answer quality than anything else on this tab, and edits refresh that image's search index automatically. Write captions for people. Write descriptions for the machine.

Choosing which image belongs in an answer

The assistant does not staple every picture to every answer, and how it decides is worth understanding.

Images are compared against the question, and relevance is judged relative to the document rather than against a fixed score. That sounds like a technicality and is not. In a deck where every slide is broadly on topic, a fixed threshold would either admit all of them or none; relative scoring still surfaces the best matches. And in a document where nothing matches, no image is attached at all, rather than the least-bad guess being presented as though it were relevant.

Two further touches. Section dividers are set aside in favor of content slides, because a title card carries the words of a topic without any of its substance and would otherwise score well. And where several images share a caption, each stays individually reachable.

What a reader gets

A relevant image appears inline with its caption in a styled band underneath, laid out like a figure in a published paper. Turn on Display description in chat and the longer description shows too.

Clicking opens a full-screen viewer, and several images in one answer become a temporary gallery you can move through without closing it. The same viewer serves the curated gallery, with a position counter, zoom from 100% to 400%, a thumbnail strip, and a downloads menu. Galleries of fifteen or more add a search box filtering by caption, description or file name.

One small piece of care worth noticing: while you are zoomed in, swiping is paused until you return to 100%, so panning around a detailed figure never accidentally flips you to the next picture.

Curating

Two library-wide toggles sit at the top of the Images tab. Include in chat answers decides whether images are offered to the assistant at all. Display description in chat decides whether readers see the long text.

With some images on and some off, the toggle reads "Mixed", tells you 12 of 40 are on, and it resolves to all-on rather than all-off. A mixed state should not be resolvable into silence by a stray click.

Per image, you can edit the caption and description in place, replace the file while keeping everything else intact, or turn off Include in chat answers for that one alone. That last control is how a decorative or sensitive picture stays browsable in the gallery while never appearing in an answer.

The Gallery Tool sub-tab controls what readers browse and in what order. Selecting images there does not by itself show the gallery, since the gallery toggle still has to be on, but emptying the gallery switches that toggle off for you. Nobody ever opens a gallery containing nothing.

Gallery Activity reports who viewed which images, when and from where, with a most-viewed ranking and CSV export. A view counts only after a reader lingers about a second, so scrolling past is not recorded as looking.

Covers

Every document can carry a cover image: the picture on its card in galleries and collections, and the banner at the top of the chat with the title over it.

Upload one, pick one from the document's own images, or generate one from a few keywords in under half a minute. Use dark overlay helps where a light title is hard to read against a busy picture. Covers display at 16:9, and an off-ratio image is center-cropped after a warning rather than silently mangled. With no cover, a document falls back to your organization's default, then to a placeholder tinted by its category.

Video works differently, and says so

Videos are not part of the trained text, and the distinction is important.

Paste a Vimeo or YouTube link, and the provider is detected, the thumbnail fetched, and the title pulled in where Vimeo supplies one. Private and unlisted Vimeo links work. Each video takes a title, an optional caption, and an optional description labeled as the detailed context used to train the assistant.

Those fields are the video's entire factual content as far as the assistant is concerned. It will point a reader to a relevant video, and it will not invent claims about what the video shows. If you want the assistant to know what is in a fifteen-minute walkthrough, that knowledge has to be in the description you wrote. This is the honest behavior, and it does mean a lazy description gets you a video nobody can find.

Clicking a card opens a full-screen player with the platform's own controls, the text along the bottom, and, with several videos, arrow-key navigation and a playlist.

When images do not appear

Almost every complaint traces to one of two causes.

Either the pictures were never extracted, because they were decorative graphics below the size threshold, or because the PDF was uploaded as the wrong type. Or the image is there but unfindable, because it has a vague caption, no description, or no index.

Captions for the reader, descriptions for the assistant. Get those two right and the rest mostly takes care of itself.