Analyze PDF file size
What Analyze PDF file size actually does
A compressor makes a file smaller without saying what it changed, and most advice about large PDFs is a list of things it might be. This reads the file's own indirect objects and reports the stored length of every stream, sorted into the categories a print or prepress report uses: images, embedded fonts, page content, form XObjects, other streams, and the structure left over. Then it names the individual objects, heaviest first, with the pixel size and the resolution each one is placed at. The command line tool prints the same figures for a scripted run over a folder.
How to use it
- Choose the PDF; its indirect objects and page operators are read inside this tab.
- Read the category budget, then the heaviest stored objects with their pixels, page and effective PPI.
- Download the JSON diagnosis, or run the CLI to get the same numbers across a directory.
Useful for
- Decide whether a 40 MB report is images, fonts or vector drawing before touching a compressor.
- Find the one scanned page holding most of a file rather than downsampling everything.
- Show a supplier which embedded object needs replacing, by object number and stored size.
Limits worth knowing
- Stored lengths are the encoded bytes of each stream, which is what compression changes; they are not the decoded pixel sizes shown separately.
- Bytes held inside compressed object streams are reported under structure and overhead rather than split among the objects within them.
- A resolution is shown for an object only when one stored image and one placement share the same pixel size; two same-size images leave that column blank rather than guessing.
- The analyzer diagnoses; it does not alter or compress the source file.
Questions people ask
Why is my PDF file so large?
Usually full-page scans or photographs stored at far more pixels than their placed size needs, and after that embedded fonts, vector drawing in page content streams, attachments, and superseded data left by incremental saves. This report answers it for the specific file rather than in general, by stored bytes per category and per object.
Are these the same numbers a desktop tool reports?
They are the same quantity: the encoded length of each stored stream, sorted into the same categories a space-audit report uses. The command line tool in this project reads the identical lengths, so a scripted run and this page agree.
Does decoded pixel memory equal PDF file size?
No, and they are kept apart here for that reason. Decoded memory is the working pixels needed to display an image; the stored stream is usually far smaller because JPEG and Flate compression are applied. Only the stored figure changes when a file is compressed.
Is the PDF uploaded?
No. Everything is read in this tab and the site has no file-upload endpoint, so there is no size cap beyond what the device can hold.