Documents: Text Extraction and Embeddings

Extract the text of a document, so that it can be split, embedded, and loaded into a vectorstore.

PDFs are read with pypdf, one page at a time, and their pages are joined with PAGE_BREAK, so that each chunk records its page. Text, Markdown, CSV, JSON and YAML are decoded as they are, and HTML has its markup removed. Other formats, e.g. Word documents, are refused: convert them to PDF first.

smarter.apps.vectorstore.extract.MAX_DOCUMENT_BYTES = 52428800

The largest document that may be added, 50 MiB.

exception smarter.apps.vectorstore.extract.VectorstoreExtractionError(message='')[source]

Bases: SmarterValueError

The document cannot be read.

smarter.apps.vectorstore.extract.content_type_for(name, content_type=None)[source]

The media type of a document, from its declared type, else from its name’s extension.

Return type:

str

smarter.apps.vectorstore.extract.extract_text(name, data, content_type=None)[source]

The text of a document.

Parameters:
  • name (str) – e.g. its file name, which determines its type if content_type does not.

  • data (bytes) – its bytes.

  • content_type (Optional[str]) – its declared media type, if any.

Return type:

tuple[str, str]

Returns:

its text, with pages separated by PAGE_BREAK, and its media type.

Raises:

VectorstoreExtractionError – if it is too large, of an unsupported type, or has no text.

The embeddings model of a vectorstore, from its Provider.

smarter.apps.vectorstore.embeddings.get_embeddings(vectorstore)[source]

The embeddings model of a vectorstore: an OpenAI-compatible embeddings API.

It uses the Provider’s base URL and API key, so any provider with an OpenAI-compatible embeddings endpoint works, including an LLMHost that serves an embeddings model, registered as a Provider.

Raises:

SmarterConfigurationError – if the vectorstore has no Provider, model or API key.

Return type:

Embeddings