Skip to content

Document & Email Summarizer ​

You will build a background service that watches a folder on your machine, detects any new document dropped in, extracts the text, runs it through a local LLM to pull out action items and key decisions, and posts a formatted summary to a Slack channel. This is for operations managers, executive assistants, and distributed teams who handle a steady stream of meeting notes, proposals, contracts, and email threads and need the signal extracted and delivered to the right people without manual forwarding.

What you will need ​

ItemDetail
StackPython 3.10+, watchdog (file watcher), Ollama (local LLM), Slack incoming webhook
Time45-60 minutes
CostFree if you run Ollama on your own hardware. Slack incoming webhooks are free. A capable local model like Llama 3.1 8B needs at minimum 8 GB of RAM.
PrerequisitesPython 3.10+ with pip or uv, Ollama installed with at least one model pulled (llama3.1:8b or mistral:7b work well), a Slack workspace where you can create an incoming webhook

How it works ​

The Python script uses the watchdog library to monitor a designated folder for new files. When a file appears, the script reads it (handling .txt, .md, .pdf, .docx), sends the text to Ollama with a summarization prompt, then posts the formatted result to a Slack channel via an incoming webhook. The script runs as a background process using systemd or a terminal multiplexer.

Build it ​

1. Install the dependencies ​

Create a project directory and install the required Python packages. The watchdog library handles filesystem events, pymupdf extracts text from PDFs, and python-docx handles Word documents.

bash
mkdir -p ~/document-summarizer
cd ~/document-summarizer
python3 -m venv venv
source venv/bin/activate
pip install watchdog pymupdf python-docx requests

If you use uv for faster installs:

bash
uv pip install watchdog pymupdf python-docx requests

2. Pull the Ollama model ​

Choose a model that fits your hardware. Llama 3.1 8B produces quality summaries on consumer hardware. Mistral 7B is faster on machines with limited RAM.

bash
ollama pull llama3.1:8b

Test that Ollama responds:

bash
ollama run llama3.1:8b "Summarize this in one sentence: The quarterly board meeting covered revenue growth of 14 percent, a new partnership with Acme Corp, and the decision to delay the European expansion until Q3."

You should see a concise one-sentence summary returned.

3. Configure the Slack webhook ​

In your Slack workspace, go to the Incoming Webhooks app page and create a new app (or use an existing one). Enable incoming webhooks, then create a webhook for the channel where summaries should appear. Copy the webhook URL. It looks like:

https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXXXXXXXXXXXXXXXXXX

Create a .env file in the project directory:

ini
SLACK_WEBHOOK_URL=https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXXXXXXXXXXXXXXXXXX
WATCH_FOLDER=/home/youruser/Documents/incoming
OLLAMA_MODEL=llama3.1:8b
OLLAMA_HOST=http://localhost:11434

4. Write the summarizer script ​

Create watcher.py. This is the complete script that ties together file watching, text extraction, Ollama summarization, and Slack posting.

python
#!/usr/bin/env python3
"""Document & Email Summarizer: watches a folder, summarizes new files, posts to Slack."""

import os
import sys
import time
import json
import logging
from pathlib import Path
from datetime import datetime

import requests
from watchdog.observers import Observer
from watchdog.events import FileSystemEventHandler

# --- Configuration from environment ---
SLACK_WEBHOOK_URL = os.environ["SLACK_WEBHOOK_URL"]
WATCH_FOLDER = os.environ.get("WATCH_FOLDER", os.path.expanduser("~/Documents/incoming"))
OLLAMA_MODEL = os.environ.get("OLLAMA_MODEL", "llama3.1:8b")
OLLAMA_HOST = os.environ.get("OLLAMA_HOST", "http://localhost:11434")

SUPPORTED_EXTENSIONS = {".txt", ".md", ".pdf", ".docx"}

logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s [%(levelname)s] %(message)s",
)
logger = logging.getLogger("docsum")


def extract_text(filepath: Path) -> str:
    """Extract text from supported file types."""
    suffix = filepath.suffix.lower()

    if suffix in {".txt", ".md"}:
        return filepath.read_text(encoding="utf-8", errors="replace")

    if suffix == ".pdf":
        import fitz  # pymupdf
        doc = fitz.open(str(filepath))
        text = "\n".join(page.get_text() for page in doc)
        doc.close()
        return text.strip()

    if suffix == ".docx":
        from docx import Document
        doc = Document(str(filepath))
        return "\n".join(p.text for p in doc.paragraphs if p.text.strip())

    raise ValueError(f"Unsupported file type: {suffix}")


def summarize_via_ollama(text: str, filename: str) -> dict:
    """Send text to Ollama and get structured summary back."""
    prompt = f"""You are a business document analyst. Read the following document and extract:

1. A one-paragraph executive summary (3-4 sentences)
2. A bullet list of action items (who needs to do what)
3. Key decisions mentioned
4. Any deadlines or dates mentioned
5. A risk flag: "HIGH", "MEDIUM", or "LOW" based on urgency or severity of content

Document filename: {filename}

Document text:
{text[:8000]}

Respond in this exact JSON format:
{{
  "summary": "...",
  "action_items": ["...", "..."],
  "key_decisions": ["...", "..."],
  "deadlines": ["...", "..."],
  "risk_flag": "LOW"
}}"""

    response = requests.post(
        f"{OLLAMA_HOST}/api/generate",
        json={
            "model": OLLAMA_MODEL,
            "prompt": prompt,
            "stream": False,
            "options": {"temperature": 0.1, "num_predict": 1024},
        },
        timeout=120,
    )
    response.raise_for_status()
    result = response.json()

    raw_output = result.get("response", "").strip()

    # Extract JSON from the response (Ollama may wrap it in markdown or extra text)
    if "```json" in raw_output:
        raw_output = raw_output.split("```json")[1].split("```")[0]
    elif "```" in raw_output:
        raw_output = raw_output.split("```")[1].split("```")[0]

    return json.loads(raw_output)


def post_to_slack(summary_data: dict, filename: str, filepath: Path):
    """Format and post the summary to Slack."""
    risk_emoji = {"HIGH": "🔴", "MEDIUM": "🟡", "LOW": "🟢"}
    emoji = risk_emoji.get(summary_data.get("risk_flag", "LOW"), "⚪")

    action_items = "\n".join(
        f"  • {item}" for item in summary_data.get("action_items", [])
    ) or "  _(none detected)_"

    decisions = "\n".join(
        f"  • {item}" for item in summary_data.get("key_decisions", [])
    ) or "  _(none detected)_"

    deadlines = "\n".join(
        f"  • {item}" for item in summary_data.get("deadlines", [])
    ) or "  _(none detected)_"

    message = {
        "blocks": [
            {
                "type": "header",
                "text": {"type": "plain_text", "text": f"{emoji} New document: {filename}"},
            },
            {
                "type": "section",
                "text": {
                    "type": "mrkdwn",
                    "text": f"*Summary:*\n{summary_data.get('summary', 'No summary generated.')}",
                },
            },
            {
                "type": "section",
                "text": {
                    "type": "mrkdwn",
                    "text": f"*Action Items:*\n{action_items}",
                },
            },
            {
                "type": "section",
                "text": {
                    "type": "mrkdwn",
                    "text": f"*Key Decisions:*\n{decisions}\n\n*Deadlines:*\n{deadlines}",
                },
            },
            {
                "type": "context",
                "elements": [
                    {
                        "type": "mrkdwn",
                        "text": f"Processed {datetime.now().strftime('%Y-%m-%d %H:%M')} | Risk: {summary_data.get('risk_flag', 'UNKNOWN')} | Source: `{filepath}`",
                    }
                ],
            },
        ]
    }

    resp = requests.post(SLACK_WEBHOOK_URL, json=message, timeout=30)
    resp.raise_for_status()
    logger.info("Posted summary to Slack for: %s", filename)


class DocumentHandler(FileSystemEventHandler):
    """Handles new file events in the watched folder."""

    def on_created(self, event):
        if event.is_directory:
            return

        filepath = Path(event.src_path)
        suffix = filepath.suffix.lower()

        if suffix not in SUPPORTED_EXTENSIONS:
            logger.info("Skipping unsupported file: %s", filepath.name)
            return

        # Wait for the file to be fully written (simple debounce)
        time.sleep(2)

        if not filepath.exists():
            return

        logger.info("Processing: %s", filepath.name)

        try:
            text = extract_text(filepath)
            if not text.strip():
                raise ValueError("No text content extracted")

            summary_data = summarize_via_ollama(text, filepath.name)
            post_to_slack(summary_data, filepath.name, filepath)
            logger.info("Completed: %s", filepath.name)

        except Exception as exc:
            logger.error("Failed to process %s: %s", filepath.name, exc)
            # Post error notification to Slack
            error_msg = {
                "text": f"⚠️ Failed to summarize `{filepath.name}`: {exc}"
            }
            try:
                requests.post(SLACK_WEBHOOK_URL, json=error_msg, timeout=10)
            except Exception:
                pass


def main():
    watch_path = Path(WATCH_FOLDER)
    watch_path.mkdir(parents=True, exist_ok=True)
    logger.info("Watching folder: %s", watch_path)
    logger.info("Model: %s | Slack webhook configured: %s",
                OLLAMA_MODEL, bool(SLACK_WEBHOOK_URL))

    handler = DocumentHandler()
    observer = Observer()
    observer.schedule(handler, str(watch_path), recursive=False)
    observer.start()

    logger.info("Document Summarizer running. Drop files into %s", watch_path)

    try:
        while True:
            time.sleep(1)
    except KeyboardInterrupt:
        observer.stop()
        logger.info("Shutting down.")
    observer.join()


if __name__ == "__main__":
    main()

5. Run the watcher ​

Make sure your environment variables are set, then start the script:

bash
source venv/bin/activate
export $(cat .env | xargs)
python watcher.py

The output confirms the watcher is active:

2026-07-15 09:23:01,142 [INFO] Watching folder: /home/youruser/Documents/incoming
2026-07-15 09:23:01,143 [INFO] Model: llama3.1:8b | Slack webhook configured: True
2026-07-15 09:23:01,144 [INFO] Document Summarizer running. Drop files into /home/youruser/Documents/incoming

Drop a .txt, .md, .pdf, or .docx file into the watched folder. Within a few seconds, a formatted summary with action items, decisions, deadlines, and a risk flag appears in your Slack channel.

6. Run as a background service ​

To keep the watcher running after you close your terminal, create a systemd user service file at ~/.config/systemd/user/docsum.service:

ini
[Unit]
Description=Document & Email Summarizer
After=network.target ollama.service

[Service]
Type=simple
WorkingDirectory=/home/youruser/document-summarizer
EnvironmentFile=/home/youruser/document-summarizer/.env
ExecStart=/home/youruser/document-summarizer/venv/bin/python /home/youruser/document-summarizer/watcher.py
Restart=on-failure
RestartSec=10

[Install]
WantedBy=default.target

Enable and start it:

bash
systemctl --user daemon-reload
systemctl --user enable docsum.service
systemctl --user start docsum.service
systemctl --user status docsum.service

To keep user services alive after logout, enable lingering:

bash
loginctl enable-linger $USER

7. Optional: Cloud fallback with Claude Code ​

If Ollama is unavailable or the local model struggles with complex documents (legal contracts, dense financial reports), you can add a fallback path that uses Claude Code via its API. Add this function to watcher.py and call it when Ollama returns an error or low-confidence result:

python
def summarize_via_claude_code(text: str, filename: str) -> dict:
    """Fallback: use Claude Code for higher-quality summarization."""
    import subprocess

    prompt = f"""You are a business document analyst. Read this document and respond with ONLY a JSON object (no markdown, no explanation). The JSON must have these keys: summary (string), action_items (array of strings), key_decisions (array of strings), deadlines (array of strings), risk_flag (one of HIGH/MEDIUM/LOW).

Document: {text[:8000]}

Respond with the JSON object only."""

    result = subprocess.run(
        ["claude", "--bare", "-p", prompt, "--max-turns", "3", "--output-format", "json"],
        capture_output=True,
        text=True,
        timeout=120,
    )

    if result.returncode != 0:
        raise RuntimeError(f"Claude Code failed: {result.stderr}")

    return json.loads(result.stdout)

This gives you a local-first pipeline with a cloud quality upgrade path when needed. You keep data local by default and only send document text to Claude's API when the local model cannot handle the task.

What goes wrong ​

MistakeSymptomFix
Ollama not runningrequests.exceptions.ConnectionError on port 11434Start Ollama with ollama serve in another terminal, or enable the systemd service: systemctl --user start ollama. Verify with curl http://localhost:11434/api/tags.
Model not pulledOllama returns 404 or "model not found"Run ollama pull llama3.1:8b (or whichever model you configured). Check available models with ollama list.
Slack webhook returns 400requests.exceptions.HTTPError: 400 Client ErrorThe Slack Block Kit format is strict. Make sure the text field inside each block exists and is a string, not None. Check for empty action_items or key_decisions lists and provide fallback text.
PDF extraction returns empty textSummary posted with "No text content extracted" errorScanned PDFs without OCR produce no extractable text. Use pymupdf with Tesseract: install pytesseract and add a fallback path that calls page.get_pixmap() then pytesseract.image_to_string(). Alternatively, pre-process scanned PDFs with OCR before dropping them in.

The result ​

You have a background service that turns your file system into a processing pipeline. To verify it works:

  1. Create a test file: echo "Meeting notes: Team agreed to launch beta on August 15. Alice will handle QA. Bob needs the API docs updated by Friday. Budget approved at $45,000." > ~/Documents/incoming/test-notes.txt
  2. Check Slack within 10 seconds. A message should appear with the summary, action items (Alice: QA, Bob: API docs), the deadline (Friday), and a risk flag
  3. Drop a real PDF into the folder and confirm the extraction and summary work
  4. Run systemctl --user status docsum.service to confirm the background service is active

The system processes any supported file the moment it lands in the folder. Team members share documents by dropping them in a shared network folder or synced cloud drive mapped to the watched directory. Summaries land in Slack where the team already works, tagged with a risk level and clear action items. No copying and pasting, no forwarding email threads, no forgotten action items from yesterday's meeting notes.