Enhance file detail page with original and processed file previews, GPT metadata display, and text extraction functionality

- Added endpoints for previewing original and processed PDF files.
- Implemented on-demand text extraction from original and processed PDFs.
- Updated file detail page to show original and processed file paths with existence status.
- Introduced GPT metadata display with a collapsible JSON view.
- Enhanced front-end with PDF.js for in-browser PDF rendering and improved user experience.
- Added integration tests for new features including metadata display and file previews.
This commit is contained in:
Christian Krakau-Louis
2026-02-11 23:32:23 +01:00
parent ce40cbcdd8
commit 02e1445e01
5 changed files with 1037 additions and 27 deletions
+111
View File
@@ -0,0 +1,111 @@
# File Detail Page - Technical Reference
## API Endpoints
```
GET /files/{id}/detail → Main detail page (template view)
GET /files/{id}/preview/original → Serve original PDF file
GET /files/{id}/preview/processed → Serve processed PDF file
GET /files/{id}/text/original → Extract text from original PDF (on-demand)
GET /files/{id}/text/processed → Extract text from processed PDF (on-demand)
POST /api/files/{id}/retry → Retry full processing
POST /api/files/{id}/retry-subtask → Retry specific task
```
## Text Extraction
Text extraction is performed **on-demand** when the user clicks "View Extracted Text":
- Uses PyPDF2 to extract text from PDF files in real-time
- Returns JSON: `{"text": "...", "page_count": 3}`
- Client-side caching prevents re-extraction on subsequent views
- Loading indicator shown during extraction
- Graceful error handling if extraction fails
## Template Context Variables
| Variable | Type | Description |
|----------|------|-------------|
| `file` | FileRecord | Database record with file metadata |
| `gpt_metadata` | dict | Extracted metadata from JSON file (or None) |
| `extracted_text` | str | Text content from OCR/extraction (or None) |
| `original_file_exists` | bool | True if original file exists on disk |
| `processed_file_exists` | bool | True if processed file exists on disk |
| `logs` | List[ProcessingLog] | Processing history records |
| `step_summary` | dict | Aggregated processing status counts |
| `flow_data` | dict | Processing flow visualization data |
## File System Structure
```
workdir/
├── tmp/
│ └── {uuid}.pdf ← Temporary ingestion file
├── original/
│ └── {uuid}.pdf ← Immutable original copy (original_file_path)
└── processed/
├── 2024-01-15_Invoice.pdf ← Processed with metadata (processed_file_path)
└── 2024-01-15_Invoice.json ← GPT metadata JSON
```
## Metadata JSON Schema
Stored as `{processed_file_path_without_extension}.json`:
```json
{
"document_type": "Invoice",
"filename": "2024-01-15_Company_Invoice",
"date": "2024-01-15",
"absender": "Test Company GmbH",
"empfaenger": "Customer Inc",
"betrag": "€ 1,234.56",
"kontonummer": "DE89370400440532013000",
"tags": ["invoice", "payment", "2024"],
"original_file_path": "/workdir/original/{uuid}.pdf",
"processed_file_path": "/workdir/processed/2024-01-15_Invoice.pdf"
}
```
## JavaScript Functions
```javascript
toggleMetadata() // Show/hide JSON metadata view
toggleTextModal(modalId) // Open/close text extraction modals
window.onclick // Close modal when clicking outside
```
## Database Schema
```sql
CREATE TABLE files (
id INTEGER PRIMARY KEY,
filehash TEXT NOT NULL,
original_filename TEXT,
local_filename TEXT NOT NULL,
original_file_path TEXT, -- Points to workdir/original/
processed_file_path TEXT, -- Points to workdir/processed/
file_size INTEGER,
mime_type TEXT,
created_at DATETIME
);
CREATE TABLE processing_logs (
id INTEGER PRIMARY KEY,
file_id INTEGER,
task_id TEXT,
step_name TEXT, -- e.g., "extract_text", "process_with_azure_document_intelligence"
status TEXT, -- "pending", "in_progress", "success", "failure"
message TEXT,
detail TEXT, -- May contain extracted text
timestamp DATETIME
);
```
## Processing Stages
Text extraction logs to check for `extracted_text`:
- `extract_text` - Local PDF text extraction
- `process_with_azure_document_intelligence` - Azure OCR processing
Both may contain extracted text in the `detail` field when `status = 'success'`.
+55 -4
View File
@@ -95,13 +95,52 @@ The **Files** page provides access to all processed documents:
When you click on a file, you'll see a comprehensive detail view with the following sections:
#### File Information
This section displays metadata about the file:
#### File Information
This section displays metadata about the file:
- File ID and original filename
- File hash (SHA-256)
- File size and MIME type
- Creation timestamp
- Local path and disk status
- Original file path and status (shows if the immutable original is available)
- Processed file path and status (shows if the final processed file is available)
#### Extracted Metadata (GPT)
If metadata has been extracted by GPT, this section displays structured information including:
- **Document Type**: Classification (Invoice, Receipt, Contract, etc.)
- **Suggested Filename**: AI-recommended filename based on content
- **Document Date**: Extracted date from document
- **Sender (Absender)**: Sender or issuing party information
- **Recipient (Empfänger)**: Recipient information
- **Amount (Betrag)**: Financial amounts (for invoices, receipts)
- **Account Number**: Bank account information
- **Tags**: Document categories and labels
**Show JSON**: Click this button to toggle the full metadata in JSON format. This is useful for:
- Viewing all extracted fields at once
- Debugging metadata extraction issues
- Copying metadata for external use
- Understanding the complete data structure
#### Document Previews
View your documents side-by-side in embedded PDF viewers:
- **Original Document**: The immutable original file as first ingested, before any processing
- **Processed Document**: The final file with embedded metadata
**Features**:
- In-browser PDF rendering for immediate viewing
- Side-by-side comparison of original vs processed versions
- Full 600px height previews for detailed review
- **View Extracted Text** buttons below each preview to see the full text content
**Text Extraction Modals**: Click "View Extracted Text" to open a fullscreen modal showing:
- Complete extracted text from OCR or PDF text layer
- Dark-themed, scrollable display for easy reading
- Copy-friendly pre-formatted text
- Close by clicking the close button or clicking outside the modal
**Note**: The original file is stored in an immutable archive (`workdir/original/`) and is never modified. The processed file is stored in `workdir/processed/` with the suggested filename and embedded metadata.
#### Processing History
View the complete processing history with a timeline showing:
@@ -138,11 +177,23 @@ If your file is still available on disk, you can preview it directly in the brow
- **Processed File**: The file after metadata has been embedded (if processing completed)
Both previews support:
- In-browser PDF viewing
- Opening in a new tab for full-screen viewing
- In-browser PDF viewing with embedded viewer
- Side-by-side comparison of original and processed versions
- Full text extraction viewing via modal overlays
**Note**: The original file is stored in an immutable archive and is never modified. This ensures you always have access to the file exactly as it was uploaded, which is valuable for:
**View Extracted Text**: Each preview includes a button to view the complete extracted text in a fullscreen modal. When you click this button:
- The system extracts text from the PDF file on-demand using PyPDF2
- A loading indicator shows while extraction is in progress
- The extracted text is displayed in a scrollable, copy-friendly format
- The text is cached so subsequent views load instantly
This is useful for:
- Verifying OCR accuracy
- Searching within document content
- Copying text for external use
- Reading documents without downloading them
**Note**: The original file is stored in an immutable archive (`workdir/original/`) and is never modified. This ensures you always have access to the file exactly as it was uploaded, which is valuable for:
- Auditing and compliance
- Debugging processing issues
- Reprocessing with improved algorithms