# Copilot Instructions for DocuElevate ## Project Overview DocuElevate is an intelligent document processing system that automates handling, extraction, and processing of documents. It integrates with multiple cloud storage providers (Dropbox, Google Drive, OneDrive, S3, Nextcloud) and uses AI services (OpenAI, Azure Document Intelligence) for metadata extraction and OCR. ## Tech Stack - **Backend**: FastAPI, SQLAlchemy, Celery, Redis - **Frontend**: Jinja2 templates, Tailwind CSS - **AI/ML**: OpenAI API, Azure Document Intelligence - **Auth**: Authentik (OAuth2), Basic Auth - **Infrastructure**: Docker, Docker Compose, Alembic (migrations) - **Testing**: Pytest, pytest-asyncio, httpx ## Supported Runtimes - **Python**: 3.11+ (3.11 and 3.12 specified in pyproject.toml) - **Docker**: Production images use `python:3.14.1` / `python:3.14.1-slim` - **Redis**: Alpine-based (`redis:alpine`) - **Gotenberg**: `gotenberg/gotenberg:latest` for PDF conversion ## Build Commands ```bash # Install production dependencies pip install -r requirements.txt # Install development dependencies (includes linters, test tools) pip install -r requirements-dev.txt # Run the FastAPI development server uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload # Run the Celery worker (requires Redis) celery -A app.celery_worker worker -B --loglevel=info -Q document_processor,default,celery # Docker build and run docker compose up -d # Database migrations alembic upgrade head # Apply all migrations alembic revision --autogenerate -m "description" # Create new migration ``` ## Test Commands ```bash # Run all tests with coverage (default via pyproject.toml addopts) pytest # Run tests by marker pytest -m unit pytest -m integration pytest -m "not requires_external" # Run a specific test file or test pytest tests/test_api.py -v pytest tests/test_api.py::test_function_name -v # Coverage report pytest --cov=app --cov-report=term-missing pytest --cov=app --cov-report=html ``` ## Lint / Format Commands ```bash # Format code with Black (line length 120) black app/ tests/ # Sort imports with isort (Black-compatible profile) isort app/ tests/ # Lint with flake8 (max line length 120, ignores E203/W503) flake8 app/ --max-line-length=120 # Type checking with mypy mypy app/ # Security lint with bandit (excludes tests) bandit -r app/ # Check for dependency vulnerabilities safety check # Run all pre-commit hooks at once pre-commit run --all-files ``` ## Core Principles ### Code Quality - Always use **Black** for formatting (line length: 120) - Use **isort** with Black profile for import sorting - Use **flake8** for linting (ignore E203, W503) - Use **type hints** for all function parameters and return values - Write **docstrings** for all public functions, classes, and modules - Maintain **80% test coverage** for new code ### Python Conventions - Use descriptive variable names (e.g., `user_document_path`, not `udp`) - Follow PEP 8 naming: `snake_case` for functions/variables, `PascalCase` for classes - Use type hints from `typing` module (Dict, List, Optional, etc.) - Prefer `pathlib.Path` over string paths for file operations - Use f-strings for string formatting, not `.format()` or `%` - Handle exceptions explicitly - avoid bare `except:` clauses ### Security Best Practices - **Never commit secrets or credentials** to the repository - Use environment variables for sensitive configuration (see `.env.demo`) - Validate and sanitize all user inputs - Use parameterized queries with SQLAlchemy (never raw SQL with user input) - Review [SECURITY_AUDIT.md](../SECURITY_AUDIT.md) before making security-related changes - Run `bandit` to check for security issues in Python code ### FastAPI Patterns - Organize endpoints by feature in `app/api/` directory - Use dependency injection for database sessions and authentication - Return Pydantic models from endpoints for automatic validation - Use proper HTTP status codes (200, 201, 400, 401, 403, 404, 500) - Document endpoints with docstrings for OpenAPI documentation - Use `async def` for I/O-bound operations ### Database (SQLAlchemy) - All models are defined in `app/models.py` - Use Alembic for schema migrations (create migration for any model change) - Use declarative base for models - Define relationships with `relationship()` and proper `back_populates` - Use database sessions from `app.database.get_db()` dependency - Always close sessions in `finally` blocks or use context managers ### Celery Tasks - Define tasks in `app/tasks/` directory, organized by feature - Use descriptive task names: `module.action` (e.g., `document.process_ocr`) - Set appropriate retry policies and error handling - Log progress and errors using Python's `logging` module - Use `bind=True` for tasks that need access to task instance - Keep tasks idempotent when possible ### Frontend - Templates are in `frontend/templates/` using Jinja2 - Static files (CSS, JS, images) in `frontend/static/` - Use Tailwind CSS utility classes (already configured) - Keep JavaScript minimal - prefer server-side rendering - Follow existing template structure and patterns ### Testing - Write tests in `tests/` directory, mirroring `app/` structure - Use pytest markers: `@pytest.mark.unit`, `@pytest.mark.integration`, etc. - Mock external services (OpenAI, Azure, cloud storage) in tests - Use `pytest.fixture` for test setup and teardown - Run tests with: `pytest -v` - Check coverage with: `pytest --cov=app --cov-report=term-missing` ### Configuration - All configuration is in `app/config.py` using Pydantic Settings - Use environment variables for configuration (12-factor app) - Provide sensible defaults when possible - Document all configuration options in `docs/ConfigurationGuide.md` ### Documentation - Keep documentation in `docs/` directory in Markdown format - Update relevant docs when adding features or changing behavior - User-facing documentation should be clear and include examples - Reference existing docs: `docs/UserGuide.md`, `docs/API.md`, `docs/DeploymentGuide.md` - See [AGENTIC_CODING.md](../AGENTIC_CODING.md) for detailed development guide ### Error Handling - Use custom exceptions defined in application (follow existing patterns) - Log errors with context using Python's `logging` module - Return user-friendly error messages in API responses - Include error details in development, sanitize in production - In API endpoints, raise `HTTPException` with appropriate status codes (400, 401, 403, 404, 500) - In Celery tasks, use `self.retry(exc=e, countdown=60)` for transient errors; log and return error dict for permanent errors - Never use bare `except:` — always catch specific exception types - Wrap database operations in `try/except` with `db.rollback()` in the except block ### Logging Conventions - Use Python's built-in `logging` module: `import logging; logger = logging.getLogger(__name__)` - **Log levels**: - `logger.debug()` — detailed diagnostic information - `logger.info()` — general operational events (document processed, task started) - `logger.warning()` — recoverable issues (retrying, fallback used) - `logger.error()` — errors that need attention (failed operations) - `logger.critical()` — system-level failures requiring immediate action - **Always include context** in log messages: `logger.info(f"Processing document: {file_id}, user: {user_id}")` - **Never log sensitive data**: passwords, tokens, API keys, personal information - Use f-strings in log messages (consistent with project style) ### Architectural Boundaries - **`app/api/`** — REST API endpoints only; organize by feature - **`app/tasks/`** — Celery background tasks only; keep idempotent - **`app/views/`** — UI routes serving Jinja2 templates - **`app/utils/`** — Shared utility functions and helpers - **`app/routes/`** — **Deprecated**; being migrated to `app/api/` — do not add new code here - **`app/models.py`** — All SQLAlchemy models (single file) - **`app/config.py`** — All configuration via Pydantic Settings (single file) - **`app/database.py`** — Database engine and session setup (single file) - **`app/auth.py`** — Authentication logic (single file) - **`frontend/templates/`** — Jinja2 templates; do not mix backend logic - **`frontend/static/`** — CSS, JS, images; keep JavaScript minimal - **`tests/`** — Test files mirroring `app/` structure - **`migrations/`** — Alembic migration scripts; always auto-generate with `alembic revision --autogenerate` ### Don't Change Rules These files and directories are managed by automation or are critical infrastructure — **do not manually edit**: - **`VERSION`** — Managed by `python-semantic-release`; updated automatically on merge to main - **`CHANGELOG.md`** — Auto-generated from conventional commit messages by semantic-release - **`migrations/`** — Do not manually edit existing migration files; only create new ones via `alembic revision --autogenerate` - **Git tags and GitHub Releases** — Created automatically by semantic-release; never create manually - **`.pre-commit-config.yaml`** — Only change if adding/updating linting tools; do not remove existing hooks - **`pyproject.toml` `[tool.semantic_release]`** — Release configuration; do not modify without explicit approval - **`docker-compose.yaml` service names** — External systems depend on `api`, `worker`, `redis`, `gotenberg` names ### Dependencies - Add new dependencies to `requirements.txt` (production) or `requirements-dev.txt` (development) - Document any new dependencies and their licenses in README.md - Check for security vulnerabilities with `safety check` - Pin major versions, allow minor updates (e.g., `fastapi>=0.100.0,<1.0.0`) ### Git Workflow - Write clear, descriptive commit messages - **ALWAYS follow Conventional Commits format** (see below) - Keep commits focused and atomic - Run tests and linters before committing - Pre-commit hooks are configured (`.pre-commit-config.yaml`) ## Conventional Commits (REQUIRED) All commit messages MUST follow the [Conventional Commits](https://www.conventionalcommits.org/) specification. ### Format ``` ():