Repository navigation
Conversation
Adds PD-layer pdDocGetAttachments and pdDocExtractAttachments, discovering files in the EmbeddedFiles name tree and FileAttachment annotations. Attachment names are sanitized, files are created exclusively and never overwritten, and streams referring to external local files are rejected. Closing a document no longer deletes files that a stream names in /F; only the parser's own temporary files are removed.
Add extraction of embedded file attachments
Owner
|
@RianKoja, thanks for the submission. I need a few days to review the submission. I will work on it mid-next week. Hope, you are ok with it. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Hi, thanks for all the good work! I'm using PDFIO.jl and would really apreciate a feature, to downlaod attachments from a pdf file. I'd be happy to contribute and iterate on revisions.
PR Content:
Adds a PD-layer API for files embedded in a PDF (implements #6):
pdDocGetAttachments,pdDocExtractAttachments(takes aPDDocor a file path),PDAttachment,pdAttachmentGetName/GetData/Extract./EmbeddedFilesname tree (including indirect/Namesand/Kids) and inFileAttachmentannotations, deduplicated by stream. Legacy/EFkeys (/Unix,/Mac,/DOS) are read too.pdDocExtractAttachments("file.pdf")writes every attachment to the current directory.Safety Measures
Names come from the PDF, so they are sanitized to a plain file name. Files are created with
O_EXCL, so nothing is overwritten or followed through a symlink; a numeric suffix is used instead. Streams pointing at a local file via/Fare ignored.One small change in shared code:
attach_object(src/CosDoc.jl) registered any stream's/Fpath for deletion on close, so a crafted PDF could make closing the document delete a local file. Only files inside the parser's own temp directory are registered now.Tests
New tests in
test/runtests.jlwith two small fixtures intest/files/: byte-exact round trip (bytes cross-checked withpypdf), name sanitizing, malformed UTF-16/UTF-8 names, collision suffixes, symlink safety, external-file rejection, a crafted PDF proving the referenced file survivespdDocClose(fails without the fix), legacy/EFkeys, and no leaked file handles. Run standalone on Julia 1.10.4: 49/49 pass. The full suite needs the externalPDTestdata, so it is left to CI.Limitations
Encrypted documents are not supported. Extraction buffers each decoded stream in memory, like the existing COS decoders.