Datasets:
Size:
1M<n<10M
ArXiv:
Tags:
Document_Understanding
Document_Packet_Splitting
Document_Comprehension
Document_Classification
Document_Recognition
Document_Segmentation
DOI:
License:
Download src/assets/models.py from amazon/doc_split: direct link, hf CLI and curl.
- Browser
- Download file 586 Bytes
-
https://huggingface.co/datasets/amazon/doc_split/resolve/refs%2Fpr%2F2/src/assets/models.py
- Command line
-
hf download hf://datasets/amazon/doc_split@refs/pr/2/src/assets/models.py
-
curl -L -o models.py https://huggingface.co/datasets/amazon/doc_split/resolve/refs%2Fpr%2F2/src/assets/models.py
586 Bytes
| # Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. | |
| # SPDX-License-Identifier: CC-BY-NC-4.0 | |
| from pydantic import BaseModel, Field | |
| class Document(BaseModel): | |
| """Represents a document in the dataset.""" | |
| doc_type: str = Field(..., description="Document category/type") | |
| doc_name: str = Field(..., description="Unique document identifier") | |
| filename: str = Field(..., description="PDF filename") | |
| absolute_filepath: str = Field(..., description="Full path to PDF file") | |
| page_count: int = Field(..., gt=0, description="Number of pages in document") | |