Method for segmenting PDF document and method for loading PDF document in webpage
Abstract
The present invention relates to a method for segmenting a PDF document and a method for loading a PDF document in a webpage. The method for segmenting a PDF document includes: S101, inspecting whether a PDF document includes an original directory structure or not; S102, if yes, segmenting the PDF document into multiple PDF sub-documents according to the original directory structure; and S103, if no, segmenting the PDF document into multiple PDF sub-documents according to document contents of the PDF document and a lexical database corresponding to the PDF document. The present invention segments a PDF document into multiple PDF sub-documents to ease operations of for example online reviewing, downloading and searching of the PDF document, and improve a user's experience of use.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A method for loading a PDF document in a webpage, comprising:
a webpage server segmenting a PDF document by using a method for segmenting a PDF document, in order to obtain multiple PDF sub-documents and PDF secondary sub-documents, and providing each of the PDF sub-documents and PDF secondary sub-documents with a corresponding network address;
the webpage server receiving a visit request transmitted from an intelligent terminal for visiting the PDF document, and the webpage server transmitting back one of the PDF sub-documents or one of the PDF secondary sub-documents of the PDF document corresponding to the visit request;
the intelligent terminal receiving and displaying the PDF sub-document or PDF secondary sub-document;
the webpage server receiving successive visit, requests transmitted from the intelligent terminal, and the webpage server transmitting back adjacently arranged PDF sub-documents or PDF secondary sub-documents among the PDF sub-documents or PDF secondary sub-documents;
the intelligent terminal receiving and displaying the PDF sub-documents or PDF secondary sub-documents; and
repeatedly executing the steps of the webpage server receiving successive visit requests transmitted from the intelligent terminal the webpage server transmitting back adjacently arranged PDF sub-documents or PDF secondary sub-documents among the PDF sub-documents or PDF secondary sub-documents and the intelligent terminal receiving and displaying the PDF sub-documents or PDF secondary sub-documents to realize continuous displaying of the PDF sub-documents and the PDF secondary sub-documents of the PDF document,
wherein the method for segmenting a PDF document comprises:
inspecting whether the PDF document includes an original directory structure or not;
if the PDF document includes an original directory structure, segmenting the PDF document into multiple PDF sub-documents according to the original directory structure; and
if the PDF document does not include an original directory structure, segmenting the PDF document into multiple PDF sub-documents according to document contents of the PDF document and a lexical database corresponding to the PDF document.
2. The method for loading a PDF document in a webpage according to claim 1 , further comprising, after the segmenting the PDF document into multiple PDF sub-document according to the original directory structure step:
determining whether or not page-number lengths of the PDF sub-documents are smaller than a first preset page-number length; and
if yes, re-segmenting the PDF document according to an upper-level directory relative to the directory level to which the PDF sub-documents correspond.
3. The method for loading a PDF document in a webpage according to claim 1 , further comprising, after the segmenting the PDF document into multiple PDF sub-document according to the original directory structure step:
determining whether or not page-number lengths of the PDF sub-documents are larger than a second preset page-number length; and
if yes, segmenting the PDF sub-documents into multiple PDF secondary sub-documents according to document contents of the PDF sub-documents and a lexical database to which the PDF sub-documents correspond.
4. The method for loading a PDF document in a webpage according to claim 3 , further comprising, after segmenting to obtain the PDF sub-documents or the PDF secondary sub-documents:
sequentially arranging the PDF sub-documents and the PDF secondary sub-documents according to positional sequences thereof in the PDF document.
5. The method for loading a PDF document in a webpage according to claim 4 , further comprising, after segmenting to obtain the PDF sub-documents or the PDF secondary sub-documents:
providing a jump tag on a document beginning or document ending of each PDF sub-document and each PDF secondary sub-document, the jump tag being provided for jumping to an adjacent PDF sub-document or PDF secondary sub-document.
6. The method for loading a PDF document in a webpage according to claim 4 , further comprising, after segmenting to obtain the PDF sub-documents or the PDF secondary sub-documents:
providing a directory tag to which each PDF sub-document and each PDF secondary sub-document corresponds, and all the directory tags forming an indexing directory for all PDF sub-documents and PDF secondary sub-documents of the PDF sub-documents.
7. The method for loading a PDF document in a webpage according to claim 6 , wherein providing a directory tag to which each PDF sub-document and each PDF secondary sub-document corresponds comprises:
generating, based on the document contents of each PDF sub-document or PDF secondary sub-document, a directory tag corresponding to each PDF sub-document or PDF secondary sub-document.
8. The method for loading a PDF document in a webpage according to claim 1 , wherein segmenting the PDF document into multiple PDF sub-documents according to document contents of the PDF document and a lexical database corresponding to the PDF document comprises:
evaluating correlation and similarity of paragraphs of the PDF sub-documents according to the document contents of the PDF sub-documents and the lexical databases to which the PDF sub-documents correspond, and taking paragraphs of which the correlation and similarity meet a preset standard as one PDF secondary sub-document.
9. The method for loading a PDF document in a webpage according to claim 1 , further comprising:
the webpage server receiving a download request transmitted from the intelligent terminal, and the webpage server transmitting back at least one of the PDF sub-documents or at least one of the PDF secondary sub-documents of the PDF document to which the download request corresponds; and
the intelligent terminal receiving and storing the PDF sub-documents or the PDF secondary sub-documents.Join the waitlist — get patent alerts
Track US11928165B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.