Implementace prevodu HTML na PDF

Sluzba prijme adresu HTML dokumentu nebo HTML v tele requestu a vrati PDF.
Navrzena pro dokumenty o stovkach az tisicich stranek.

Rendering:
- WeasyPrint jako vychozi engine, spravne CSS Paged Media, nizka pametova
  narocnost, bez JavaScriptu
- Chromium pres Playwright pro dokumenty dokreslovane skripty
- rezim auto s detekci skriptu a fallbackem pri selhani WeasyPrintu

Velke dokumenty:
- deleni na casti na strukturalnich hranicich, rez nikdy uvnitr tabulky
  nebo odstavce
- dvoupruchodovy render obsahu se skutecnymi cisly stranek, pozice nadpisu
  se ctou z kotev hlasenych u kazde stranky
- cislovani stranek bud pres CSS countery, nebo pres cislovaci vrstvu
  nastampovanou na hotove PDF, rozmer stranky se cte z vysledneho souboru
- Chromium se restartuje po N jobech, nikdy vsak behem beziciho renderu

API:
- POST /convert synchronne, POST /jobs asynchronne se sledovanim stavu,
  stahovanim vysledku, rusenim a volitelnym callbackem
- GET /health s overenim dostupnosti obou enginu a stavem fronty
- OpenAPI respektuje prefix reverse proxy pres root_path

Bezpecnost a provoz:
- SSRF kontrola po DNS resolvu, na kazdem presmerovani a u vsech pozadavku
  prohlizece
- nedostupne assety render nezastavi, ale hlasi se v odpovedi i v logu
- fronta s omezenym poctem workeru, rozpracovane joby se pri ukonceni
  oznaci jako failed, nezmizi potichu
- strukturovane JSON logovani s job_id
- vsechny limity vypnute ve vychozim stavu

Dockerfile je dvoufazovy, obsahuje zavislosti WeasyPrintu, Chromium
a fonty s ceskou diakritikou.

Autentizace zamerne neni implementovana, zpusob predavani neni domluveny.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
JiriUhlir
2026-08-27 14:50:10 +02:00
co-authored by Claude Opus 5
parent e3cc8f418b
commit 156289fe2d
48 changed files with 4043 additions and 24 deletions
+80
View File
@@ -0,0 +1,80 @@
"""Page numbering of the merged document.
Two ways to number pages:
css
CSS counters do the work during the render. Correct and cheap, but it only
works when the whole document is rendered in one pass, because each chunk
restarts the page counter.
overlay
A transparent numbering layer with the same page size is rendered once and
merged onto the finished PDF. This is the only option for a chunked document
and for Chromium, which has no usable page counters.
"""
from __future__ import annotations
import logging
from pathlib import Path
from ..errors import EngineUnavailableError
from ..models import PageNumbers, PageSettings
from .styles import build_overlay_css
logger = logging.getLogger(__name__)
def build_overlay(
total_pages: int,
width_pt: float,
height_pt: float,
page: PageSettings,
page_numbers: PageNumbers,
output_path: Path,
) -> Path:
"""Render the numbering layer, one empty page per page of the document."""
try:
from weasyprint import HTML
except ImportError as exc:
raise EngineUnavailableError(
"Cislovani stranek vyzaduje nainstalovany WeasyPrint, ktery kresli cislovaci vrstvu.",
) from exc
css = build_overlay_css(width_pt, height_pt, page, page_numbers, total_pages)
slots = '<div class="pdf-page-slot"></div>' * total_pages
html = (
"<!DOCTYPE html><html><head><meta charset=\"utf-8\">"
f"<style>{css}</style></head><body>{slots}</body></html>"
)
HTML(string=html).write_pdf(target=str(output_path))
logger.info("Numbering overlay rendered", extra={"pages": total_pages})
return output_path
def apply_overlay(document_path: Path, overlay_path: Path, output_path: Path) -> None:
"""Stamp the numbering layer onto every page of the document."""
from pypdf import PdfReader, PdfWriter
writer = PdfWriter(clone_from=str(document_path))
try:
with overlay_path.open("rb") as handle:
overlay = PdfReader(handle)
available = len(overlay.pages)
if available < len(writer.pages):
logger.warning(
"Numbering overlay has fewer pages than the document, tail will stay unnumbered",
extra={"overlay_pages": available, "document_pages": len(writer.pages)},
)
for index, page in enumerate(writer.pages):
if index >= available:
break
page.merge_page(overlay.pages[index])
with output_path.open("wb") as target:
writer.write(target)
finally:
writer.close()