Skip to content

How to extract feature vectors from a single PDF file for inference? #18

Description

@Raneem04H

Hello,

Thank you for releasing the EMBER2024 dataset and the thrember package.

I trained a CatBoost classifier on the EMBER2024 PDF subset using the vectorized features.

Now I would like to deploy the model and classify new PDF files.

I found PEFeatureExtractor for PE files, but I could not find an equivalent feature extractor for PDF files.

Is there an official way to extract EMBER2024 feature vectors from a single PDF file for inference?

If such functionality exists but is not currently exposed in the Python package, could you please point me to the relevant code or example?

Thank you very much.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions