Dev.to · 4 min read

RAG Database with VLM as Extractor

RAG Database with VLM as Extractor

Problem Building an effective semantic retrieval pipeline is challenging when the input data contains inconsistent formats, including: Images Random spreadsheets PDFs and documents Mixed structured and unstructured attachments Traditional chunking methods struggle because the content structure varies significantly between files. Extracting meaningful text representations before embedding is required to improve retrieval quality. Solution: VLM-Based Content Extraction To handle heterogeneous input formats, the pipeline was redesigned to use a Vision Language Model (VLM) as an extraction layer. Instead of relying on traditional parsers, each document page is converted into an image and processed independently by the VLM. Pipeline Overview Upload document attachment Extract pages and convert each page into an image Encode images into Base64 format Send images to VLM for summarization and keyword extraction Embed the generated summaries Store embeddings in PostgreSQL with pgvector Use semantic search for retrieval by the AI agent Image Conversion Pipeline Each page is treated as an independent chunk. async def vlm_inference(self, vlm_req: VLMRequest, file: UploadFile) -> list[str]: img_list = await extract_content(file, True) results = [] for idx, img in enumerate(img_list): # Save image for inspection image_path = save_dir / f"page_{idx}.png" img.save(image_path) print(f"Saved image: {image_path}") # Convert to base64 for VLM img_base64 = image_to_base64(img) client = OllamaClient() result = await client.vlm_inference( vlm_req=vlm_req, image_base64=img_base64, ) results.append(result) return results VLM Client post: async def vlm_inference( self, vlm_req: VLMRequest, image_base64: str ) -> str: payload = { "model": vlm_req.model, "messages": [ { "role": "user", "content": vlm_req.prompt, "images": [image_base64], } ], "stream": False, "think": vlm_req.thinking, } response = requests.post( f"{LLM_BASE_URL}/api/chat", json=payload, timeout=300 ) response.raise_for_status() print(response.json()) return response.json()['message']['content'] The advantage of page-level chunking is that it avoids complex document-specific parsing logic while preserving the visual context of the original file. VLM Inference Client The extracted page image is sent to the VLM through the Ollama API. async def vlm_inference( self, vlm_req: VLMRequest, image_base64: str ) -> str: payload = { "model": vlm_req.model, "messages": [ { "role": "user", "content": vlm_req.prompt, "images": [image_base64], } ], "stream": False, "think": vlm_req.thinking, } response = requests.post( f"{LLM_BASE_URL}/api/chat", json=payload, timeout=300 ) response.raise_for_status() return response.json()['message']['content'] Model Selection and Optimization Initially, the extraction model used was: Qwen3.5:20B However, inference latency was too high for processing thousands of attachments. The pipeline was optimized by switching to: Qwen3.5:4B The smaller model provided a better balance between: Processing speed Resource consumption Extraction quality ## Thinking Mode Optimization The Qwen model's reasoning mode introduced a practical issue. With thinking enabled: The model spent additional tokens on internal reasoning Long generation consumed the output token budget Some pages returned empty summaries because the model stopped after reaching the maximum output length Increasing the token limit was not desirable because: The system processes thousands of attachments Larger outputs increase inference cost Storage requirements increase significantly Therefore, the reasoning mode was disabled. With thinking disabled: Response latency improved Output became more concise Important keywords were still preserved Retrieval performance remained acceptable Embedding and Storage After VLM extraction, the generated summaries are converted into embeddings. The embeddings are stored in PostgreSQL using pgvector. The VLM-based extraction approach significantly improved semantic retrieval performance by creating consistent text representations from heterogeneous documents. Result The AI agent can now retrieve relevant information from: Images Scanned documents Tables Mixed-format attachments without requiring individual parsers for every file type. The key design decision was using the VLM as a universal extraction layer, transforming unstructured visual information into searchable semantic representations.

This is a summary aggregated from Dev.to. Read the complete article on the original site:

Read full article at Dev.to

More AI & Machine Learning News