Base64 carries bytes and nothing else: no file name, no type, not even a sign that it is complete. So the decoder reads the decoded bytes to say what the file is, from the signatures in the WHATWG MIME Sniffing standard and those of other common formats. A PDF starts with %PDF- and its version, which in Base64 is always JVBERi0; a ZIP starts with PK, which is UEsDB. Formats built on ZIP are told apart by the names in the archive’s central directory, read without decompressing anything: [Content_Types].xml with a word/ folder is a .docx, a stored mimetype entry names an OpenDocument file or an EPUB, and AndroidManifest.xml makes an APK. Legacy .doc, .xls and .msg files are told apart by their OLE2 directory. An API that labels everything application/octet-stream, or a data: URI that claims image/png, still gets a download with the right extension, and the disagreement is pointed out.
For a PDF, the header gives the version, and a later /Version in the document catalog overrides it. The page count is the page tree’s own /Count. PDF 1.5 and later can keep that inside compressed object streams, so those are inflated to find it; if the page tree still cannot be found, the /Type /Page objects are counted and the number is marked approximate, and when the object streams are encrypted the count is reported as unknown. An /Encrypt entry in the trailer means the file is encrypted: the algorithm and the permissions it denies are read from the encryption dictionary, but whether a password is needed to open it cannot be told without trying one. A PDF with no %%EOF was cut off; JavaScript and attached files are flagged. The download is exactly the bytes decoded — nothing is repaired or decrypted — and nothing is previewed on the page; HTML and SVG are never rendered.
Base64 is accepted as it arrives: wrapped at 64 or 76 columns, in the URL-safe alphabet, without padding, as a JSON string with its quotes and \n escapes, or as a whole JSON response, where the longest Base64-looking value is used and a filename or contentType field beside it is honoured. A data: URI of any type, an email attachment pasted with its MIME headers, and a PEM block work too. A length one character past a group of four cannot come from any encoder, so it is reported as cut off; decoded bytes that are themselves Base64 text were encoded twice, and are offered for decoding again. Encoding gives the Base64 on one line or wrapped for PEM or MIME, a data URI, which is never wrapped, and a JSON string. Files up to 3.75 MB (5 MB as Base64) are accepted; anything over about 750 KB is processed in a background worker, so the page does not freeze.