
AS
Albert Sawczyn, Jakub Binkowski, Kamil Tagowski, {\L}ukasz Augustyniak, Berenika Kaczmarek-Templin, Tomasz Kajdanowicz
· 1 min read
ResearcharXiv cs.CL
Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research
arXiv:2609.22209v1 Announce Type: new
Abstract: Empirical legal research often relies on turning research questions into structured data extracted from large collections of rulings and judgments. Designing the extraction schema and then extracting the data remain a manual, expertise-heavy bottleneck. We present schematize, an open-source multi-agent system that interactively turns a researcher's problem statement into a validated extraction schema that can later be used for autonomous extraction. Schematize couples (i) a clarification dialogue that elicits implicit expert intent, (ii) iterative schema generation, (iii) data-grounded refinement that tests the schema against documents, and (iv) chat-based post-editing. We evaluated the system with human legal professional, introducing our novel methodology, and schematize achieves top performance in most of tested configurations. While the system is designed to be domain-agnostic and applicable to any document collection, we tailor and evaluate it on legal research problems. We release schematize as a pip-installable Python package with full documentation.
Original source
This story was published by arXiv cs.CL and written by Albert Sawczyn, Jakub Binkowski, Kamil Tagowski, {\L}ukasz Augustyniak, Berenika Kaczmarek-Templin, Tomasz Kajdanowicz. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


