Status: Draft (as of September 2026). There is no compatibility guarantee yet.
SUITAR is a common-denominator, trivial-to-parse interchange format for “the result of decompressing a compressed file or unpacking an archive file”, such as ZIP, RAR, 7Z, TAR, GZIP, BZIP2, XZ, ZSTD, etc.
SUITAR is to archive files (including foobar.dat.xz compressed files as an archive containing 1 file) as Farbfeld or NIE is to image files: an uncompressed, “designed for Unix pipes” format that is trivial to read or write in a few hundred lines of code, ideally in a memory-safe programming language. It‘s a format for what the Chromium web browser’s [https://chromium.googlesource.com/chromium/src/+/master/docs/security/rule-of-2.md](Rule of 2) security advice calls Normalization.
SUITAR is a subset of the well-known and widely-used TAR (Tape Archive, with GNU extensions) archive file format. Every valid SUITAR file is also a valid TAR file. These files work with popular tools like /usr/bin/tar and with popular TAR-reading libraries in a variety of programming languages.
SUITAR is a subset of “pure TAR”, not of “TAR wrapped in GZIP”, “TAR wrapped in BZIP2”, etc. A separate compression step wrapping SUITAR is feasible, just like wrapping TAR, but is out of scope of this document.
Like TAR, SUITAR files are a sequence of independent entries (files or directories). Independent means that, like JSON map keys, duplicate names are valid, although some SUITAR decoders (or some destination File Systems or Operating Systems, if unpacking from a source SUITAR archive to a particular destination FS / OS) may choose to reject them.
For example, answering “are these two file names duplicates (and does the destination FS / OS accept or reject duplicates)?” may depend on the destination FS / OS's case-sensitivity and Unicode normalization configuration (and whether case folding and normalization elides DICPs, Unicode Default-Ignorable Code Points), not the source SUITAR archive per se.
A file‘s entry does not need to be preceded by explicit entries for that file’s parent directories.
Compared to plain TAR (and refer to the GNU TAR manual), SUITAR has further restrictions:
REGTYPE or AREGTYPE), sparse files (GNUTYPE_SPARSE) or directories (DIRTYPE). There is no support for hard links, symlinks, device files or other non-standard files.GNUTYPE_LONGNAME header block, regardless of whether the name's length is over or under 100 or 255 bytes.(1 << 53), which is 9007_199254_740992.rw-r--r--) or 0o755 (rwxr-xr-x), encoded in base-8 octal (not base-256).There is no support for various TAR variants, such as “the PAX extensions to TAR” or “the USTAR extensions to TAR”, other than what's implied by the subset of the GNU extensions that SUITAR explicitly uses.
These rules apply to both file names and directory names.
'\n' or '\x00', the “new line” or NUL bytes.'\x7F' ASCII DEL byte."", "." or ".."."/", "./" or "../"."/", "/." or "/.."."//", "/./" or "/../" as substrings.For example, when converting from ZIP (with Japanese file names) to SUITAR, it is the SUITAR producer‘s responsibility, not the SUITAR consumer’s, to detect and transform Shift-JIS encoded names to equivalent and valid UTF-8.
Names like "abc/.\u200B./xyz", containing a DICP (U+200B ZERO WIDTH SPACE), are valid SUITAR file names per se. However, when unpacking from a source SUITAR archive to a particular destination FS / OS, SUITAR decoders may wish to impose further restrictions on file names, to avoid path traversal attacks specific to that FS / OS. For example, eliding DICPs before applying the File Name Validity rules would prohibit the "/.\u200B./" substring.
SUITAR files are a sequence of entries, followed by a 1024-byte “End Of File” marker. The EOF marker's bytes are all NUL. Each entry occupies an integer number of 512-byte blocks:
GNUTYPE_LONGNAME header block.REGTYPE, AREGTYPE, GNUTYPE_SPARSE or DIRTYPE header block.REGTYPE, 0 or more payload blocks containing the file contents.AREGTYPE, see “Implicitly-Sized Files” below.REGTYPE or AREGTYPE, no further blocks.For TAR itself, each entry‘s contents is preceded by the entry’s header that states the entire contents' size in bytes. SUITAR keeps that formal structure but also uses additional convention to represent entries whose size is only known at the end, not the start, of the entry.
For example, when decompressing foobar.dat.gz from a GZIP stream to SUITAR (an archive with one entry: foobar.dat), the uncompressed foobar.dat size is not known until the end of the GZIP-formatted input stream is reached.
For historical reasons, TAR itself has two typeflag codes for regular files (REGTYPE and AREGTYPE) and both codes are largely equivalent. SUITAR, by convention, treats them differently: REGTYPE is for the common case, where the contents' size is known up-front, and AREGTYPE means the contents that follow are partial and to be continued.
An implicitly-sized file is partitioned into 2 or more chunks. The final chunk is REGTYPE and all other chunks are AREGTYPE. Each chunk‘s size is explicit (and zero is a valid size) but the number of chunks isn’t known until the final REGTYPE chunk is delivered. The entry's structure is:
GNUTYPE_LONGNAME header block.AREGTYPE header block, stating the chunk contents' size.REGTYPE header block, stating the chunk contents' size.“0 non-final chunks” is actually valid in some sense, equivalent to an explicitly-sized REGTYPE entry (of exactly one chunk).
The file name still comes from the GNUTYPE_LONGNAME payload. There is only one GNUTYPE_LONGNAME header block, not one per chunk.
Other metadata (mode and modTime, but not file size) comes from the initial chunk's header block. For non-initial chunks, the mode must be "644" and the modTime must be zero.
Other TAR-reading tools and libraries, which do not understand SUITAR's “implicitly-sized files” convention, will fall back to treating all non-initial chunks as separate regular files. These will all have the same file name ("\x13sUItAR", due to the SUITAR magic signature), which does not satisfy the File Name Validity rules, but this fallback name will not be presented by decoders that understand the convention.
SUITAR encoders are discouraged from using this implicitly-sized files convention unless the surrounding context ensures that the decoders also speak SUITAR, not just TAR per se. This can be more likely if writing SUITAR over a Unix pipe (with a known program on the other end), compared to writing to disk.
Like all blocks, each header block is 512 bytes long. Each header block also starts with a 12-byte magic signature (that is not valid UTF-8), identifying SUITAR version 1. There are no other versions at this time.
The first 384 out of 512 bytes must match this template (arranged as 24 rows of 16 bytes per row, plus commentary):
@@@@@@@@@@@@.... @@@@@@@@@@@@ = "\x13sUItAR\x00\xFE\xFDv1". ................ ................ ................ ................ ................ ....0000???.0177 ??? = mode. 776.0177776.$... ????????$...???? ???????? = physical size, ???????? = modTime. ??????????. ?... ?????? = checksum, ? = type. ................ ................ ................ ................ ................ ................ .ustar .nobody. ................ .........nobody. ................ ................ ................ ................ ................
In this template, @ indicates the magic signature, . indicates a 0x00 NUL byte, $ indicates a 0x80 byte and ? indicates parts of the template that are variable, not hard-coded.
These ? bytes are the 3-byte mode ("644" or "755"), physical size or modTime as an 8-byte big-endian uint64, 6-byte checksum (see below) or 1-byte type, which must be one of:
'\x00' for AREGTYPE.'0' for REGTYPE.'5' for DIRTYPE, in which case mode must be "755" and physical size must be all zeroes.'L' for GNUTYPE_LONGNAME, in which case mode must be "644" and modTime must be all zeroes.'S' for GNUTYPE_SPARSE, in which case physical size must be all zeroes and offset and logical size (see below) must be the same number and, again, within the half-open range 0 .. (1 << 53).The last 128 out of 512 bytes (8 rows of 16 bytes per row) must be all NUL bytes unless the type is GNUTYPE_SPARSE, in which case it must match this template (and ? again indicates an 8-byte big-endian uint64):
..$...????????$. ???????? = sparse offset. ................ ................ ................ ................ ................ ...$...????????. ???????? = sparse logical size. ................
A 512-byte header block's checksum value is simply 256 plus the sum of each byte (after converting from uint8 to uint32, to avoid overflow) in the block, at offsets in the two half-open ranges 0 .. 148 and 156 .. 512, which excludes the 8 bytes for the 6-byte checksum itself plus another two hard-coded bytes "\x00\x20".
That “256 plus” is equivalent to summing over the entire 0 .. 512 range if valuing the 8 checksum bytes in the range 148 .. 156 as being '\x20'.
That checksum value is written as a 6-byte ASCII octal number in the header. For example, 4853 (decimal) would be encoded as "011365" (octal).
Each entry has one or more payload blocks, between its two header blocks, containing the file or directory name. The name length (including a trailing NUL byte) is the first header block‘s physical size value, and must be within the half-open range 2 .. 4096, and so the excluding-a-trailing-NUL length must range within 1 .. 4095. Rounding up that including-a-trailing-NUL length to a multiple of 512 gives the number of 512-byte payload blocks that contain the name. All padding bytes in the name’s final payload block must be NUL.
For REGTYPE or AREGTYPE entries, the name payload is followed by one or more chunks. Concatenating the chunks' contents reconstructs the file. Only the final chunk is REGTYPE and all others are AREGTYPE. Each chunk has one header block and zero or more payload blocks. The header block‘s physical size value gives the chunk’s size and rounding that up to a multiple of 512 gives the number of 512-byte payload blocks that contain the chunk contents. Again, all padding bytes in the contents' final payload block must be NUL.
For other entries (DIRTYPE or GNUTYPE_SPARSE), there are no further payload blocks after the second header block.
For GNUTYPE_SPARSE entries, the second header block‘s logical size value gives the reconstructed file’s size and its contents are all NUL bytes.
The google/wuffs repository, which holds this specification document, also holds a suitar Go package and some test/data/*.suitar example files, readable by that Go package but also by /usr/bin/tar.
Updated on September 2026.