30-SECOND SUMMARY
What to take away
- Bad extraction and retrieval cannot be repaired reliably by a larger generator.
- Use the same embedding model for indexing and querying.
- Evaluate retrieval and answer generation separately with questions whose source locations are known.
From document to grounded answer
Retrieval and generation fail in different ways.
- 01SPLIT
Meaningful chunks
- 02EMBED
One embedding model
- 03SEARCH
Retrieve evidence
- 04ANSWER
Cite the source
The RAG pipeline
Extract text, split it into meaningful chunks, create embeddings, retrieve the closest chunks for a question, and place those chunks in the generation prompt.
Store file, page, heading, and chunk identifiers so every answer can link back to evidence.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Fix extraction before retrieval
PDF reading order, tables, headers, and scanned pages often produce broken text. Compare extracted samples with the original before indexing thousands of pages.
Remove duplicates and obsolete versions, and keep documents with different access permissions in separate indexes.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Chunk by meaning
Very short chunks lose context; very long chunks add noise. Prefer headings and paragraph boundaries, then use overlap only where continuity is genuinely needed.
Each known answer should be traceable to a specific chunk ID.
Define completion with an observable result instead of a general impression. Repeat the same input, and if the output changes, isolate whether the model, runtime settings, or source data changed before moving to the next stage.
Create and search embeddings
Ollama exposes `/api/embed` for local embeddings. Use the same embedding model for documents and questions and compare vectors with an appropriate similarity measure.
Inspect top results manually before adding filters, hybrid search, or reranking.
After running the command or code, inspect the exit status, logs, and the file, process, or response it was meant to create. If it fails, change one input, version, permission, or resource condition at a time and repeat the same check so that the cause remains attributable.
curl http://localhost:11434/api/embed -d '{
"model": "embeddinggemma",
"input": ["Refunds are available within seven days."]
}'Evaluate retrieval and generation
Create questions for single facts, multiple sources, similar documents, and answers that do not exist. First check whether the correct chunk was retrieved; then check whether the model used it accurately.
When deleting a document, remove source text, chunks, embeddings, caches, and related logs, then confirm that it can no longer be retrieved.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Frequently asked questions
Is RAG the same as training?
No. Typical RAG leaves model weights unchanged and supplies retrieved context at request time. For a practical check, follow the “The RAG pipeline” section, change one condition at a time, and record the result.
Must the embedding and chat model be the same?
No, but document indexing and query embedding should use the same embedding model. Use the same embedding model for indexing and querying. For a practical check, follow the “Fix extraction before retrieval” section, change one condition at a time, and record the result.
Does adding more documents always help?
No. Duplicates, obsolete versions, and mixed permissions can reduce quality and create security problems. For a practical check, follow the “Chunk by meaning” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Ollama Embeddings Ollama API Introduction