Sluzba prijme adresu HTML dokumentu nebo HTML v tele requestu a vrati PDF. Navrzena pro dokumenty o stovkach az tisicich stranek. Rendering: - WeasyPrint jako vychozi engine, spravne CSS Paged Media, nizka pametova narocnost, bez JavaScriptu - Chromium pres Playwright pro dokumenty dokreslovane skripty - rezim auto s detekci skriptu a fallbackem pri selhani WeasyPrintu Velke dokumenty: - deleni na casti na strukturalnich hranicich, rez nikdy uvnitr tabulky nebo odstavce - dvoupruchodovy render obsahu se skutecnymi cisly stranek, pozice nadpisu se ctou z kotev hlasenych u kazde stranky - cislovani stranek bud pres CSS countery, nebo pres cislovaci vrstvu nastampovanou na hotove PDF, rozmer stranky se cte z vysledneho souboru - Chromium se restartuje po N jobech, nikdy vsak behem beziciho renderu API: - POST /convert synchronne, POST /jobs asynchronne se sledovanim stavu, stahovanim vysledku, rusenim a volitelnym callbackem - GET /health s overenim dostupnosti obou enginu a stavem fronty - OpenAPI respektuje prefix reverse proxy pres root_path Bezpecnost a provoz: - SSRF kontrola po DNS resolvu, na kazdem presmerovani a u vsech pozadavku prohlizece - nedostupne assety render nezastavi, ale hlasi se v odpovedi i v logu - fronta s omezenym poctem workeru, rozpracovane joby se pri ukonceni oznaci jako failed, nezmizi potichu - strukturovane JSON logovani s job_id - vsechny limity vypnute ve vychozim stavu Dockerfile je dvoufazovy, obsahuje zavislosti WeasyPrintu, Chromium a fonty s ceskou diakritikou. Autentizace zamerne neni implementovana, zpusob predavani neni domluveny. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
47 lines
1.5 KiB
Python
47 lines
1.5 KiB
Python
"""Shared fixtures.
|
|
|
|
Nothing here starts a real browser. Engine level tests skip themselves when the
|
|
engine is not installed in the current environment.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import pytest
|
|
|
|
|
|
def build_document(sections: int, paragraphs_per_section: int = 12) -> str:
|
|
"""Synthetic document with predictable structure and size."""
|
|
parts = [
|
|
"<!DOCTYPE html><html><head><meta charset='utf-8'>",
|
|
"<style>body{font-family:sans-serif;font-size:11pt}</style>",
|
|
"</head><body>",
|
|
]
|
|
for index in range(sections):
|
|
parts.append(f"<section id='sekce-{index}'><h1>Kapitola {index + 1}</h1>")
|
|
for paragraph in range(paragraphs_per_section):
|
|
parts.append(
|
|
f"<p>Odstavec {paragraph + 1} kapitoly {index + 1}. "
|
|
+ ("Text s ceskou diakritikou pro overeni fontu. " * 8)
|
|
+ "</p>"
|
|
)
|
|
parts.append("</section>")
|
|
parts.append("</body></html>")
|
|
return "".join(parts)
|
|
|
|
|
|
@pytest.fixture
|
|
def small_document() -> str:
|
|
return build_document(sections=3, paragraphs_per_section=2)
|
|
|
|
|
|
@pytest.fixture
|
|
def large_document() -> str:
|
|
"""Roughly twelve hundred pages, used for memory and timing measurements."""
|
|
return build_document(sections=400, paragraphs_per_section=14)
|
|
|
|
|
|
@pytest.fixture
|
|
def unsplittable_document() -> str:
|
|
body = "".join(f"<p>Odstavec {index}. " + ("Text. " * 40) + "</p>" for index in range(200))
|
|
return f"<!DOCTYPE html><html><head><meta charset='utf-8'></head><body>{body}</body></html>"
|