Documents: Text Extraction and Embeddings
Extract the text of a document, so that it can be split, embedded, and loaded into a vectorstore.
PDFs are read with pypdf, one page at a time, and their pages
are joined with PAGE_BREAK, so that each chunk records
its page. Text, Markdown, CSV, JSON and YAML are decoded as they are, and HTML has its markup
removed. Other formats, e.g. Word documents, are refused: convert them to PDF first.
- smarter.apps.vectorstore.extract.MAX_DOCUMENT_BYTES = 52428800
The largest document that may be added, 50 MiB.
- exception smarter.apps.vectorstore.extract.VectorstoreExtractionError(message='')[source]
Bases:
SmarterValueErrorThe document cannot be read.
- smarter.apps.vectorstore.extract.content_type_for(name, content_type=None)[source]
The media type of a document, from its declared type, else from its name’s extension.
- Return type:
- smarter.apps.vectorstore.extract.extract_text(name, data, content_type=None)[source]
The text of a document.
- Parameters:
- Return type:
- Returns:
its text, with pages separated by PAGE_BREAK, and its media type.
- Raises:
VectorstoreExtractionError – if it is too large, of an unsupported type, or has no text.
The embeddings model of a vectorstore, from its Provider.
- smarter.apps.vectorstore.embeddings.get_embeddings(vectorstore)[source]
The embeddings model of a vectorstore: an OpenAI-compatible embeddings API.
It uses the Provider’s base URL and API key, so any provider with an OpenAI-compatible embeddings endpoint works, including an LLMHost that serves an embeddings model, registered as a Provider.
- Raises:
SmarterConfigurationError – if the vectorstore has no Provider, model or API key.
- Return type:
Embeddings