27c5d86527532a8a2d73ae0e6d78bdbd626c3590
This is a performance-tuning change that shouldn't change the behavior of the library. This adds some complexity but the performance gain might make that worthwhile: With FDO on perflab/haswell, a 4.0% gain (geometric mean). SAMPLE (before) Benchmark Time(ns) CPU(ns) Iterations ------------------------------------------------ BM_UFlat/0 36638 36552 100000 2.6GB/s html BM_UFlat/1 457153 455895 9173 1.4GB/s urls BM_UFlat/2 5850 5837 685481 19.6GB/s jpg BM_UFlat/3 122 122 34551988 1.5GB/s jpg_200 BM_UFlat/4 6797 6781 620811 14.1GB/s pdf BM_UFlat/5 179485 179037 23471 2.1GB/s html4 BM_UFlat/6 142734 142384 29525 1018.7MB/s txt1 BM_UFlat/7 125233 124924 33709 955.6MB/s txt2 BM_UFlat/8 382548 381533 10000 1066.7MB/s txt3 BM_UFlat/9 525614 524297 8018 876.5MB/s txt4 BM_UFlat/10 34946 34868 100000 3.2GB/s pb BM_UFlat/11 149548 149208 28063 1.2GB/s gaviota BM_UFlat/12 10684 10663 392580 2.1GB/s cp BM_UFlat/13 5494 5484 766584 1.9GB/s c BM_UFlat/14 1691 1688 2488784 2.1GB/s lsp BM_UFlat/15 676443 674726 6129 1.4GB/s xls BM_UFlat/16 156 156 26656909 1.2GB/s xls_200 BM_UFlat/17 239911 239297 17558 2.0GB/s bin BM_UFlat/18 182 182 23072932 1047.9MB/s bin_200 BM_UFlat/19 21544 21499 194484 1.7GB/s sum BM_UFlat/20 2236 2232 1877810 1.8GB/s man BM_UFlatSink/0 42266 42179 99732 2.3GB/s html BM_UFlatSink/1 461810 460633 9055 1.4GB/s urls BM_UFlatSink/2 5816 5804 632829 19.8GB/s jpg BM_UFlatSink/3 124 123 34351698 1.5GB/s jpg_200 BM_UFlatSink/4 7173 7157 609929 13.3GB/s pdf BM_UFlatSink/5 184795 184302 22660 2.1GB/s html4 BM_UFlatSink/6 143552 143223 29272 1012.7MB/s txt1 BM_UFlatSink/7 127160 126890 33178 940.8MB/s txt2 BM_UFlatSink/8 382219 381313 10000 1067.3MB/s txt3 BM_UFlatSink/9 528042 526713 7988 872.5MB/s txt4 BM_UFlatSink/10 41389 41305 100000 2.7GB/s pb BM_UFlatSink/11 147215 146877 28854 1.2GB/s gaviota BM_UFlatSink/12 12008 11984 348139 1.9GB/s cp BM_UFlatSink/13 5444 5433 775084 1.9GB/s c BM_UFlatSink/14 1647 1644 2552119 2.1GB/s lsp BM_UFlatSink/15 665011 663424 6320 1.4GB/s xls BM_UFlatSink/16 153 153 27571837 1.2GB/s xls_200 BM_UFlatSink/17 239735 239169 17411 2.0GB/s bin BM_UFlatSink/18 183 182 23005573 1046.8MB/s bin_200 BM_UFlatSink/19 22544 22498 187705 1.6GB/s sum BM_UFlatSink/20 2190 2186 1917894 1.8GB/s man SAMPLE (after) Benchmark Time(ns) CPU(ns) Iterations ------------------------------------------------ BM_UFlat/0 33940 33889 100000 2.8GB/s html BM_UFlat/1 440728 439944 9586 1.5GB/s urls BM_UFlat/2 5652 5641 744776 20.3GB/s jpg BM_UFlat/3 123 123 34647884 1.5GB/s jpg_200 BM_UFlat/4 6628 6615 631892 14.4GB/s pdf BM_UFlat/5 169523 169227 24197 2.3GB/s html4 BM_UFlat/6 144139 143892 29232 1008.0MB/s txt1 BM_UFlat/7 127148 126915 33144 940.6MB/s txt2 BM_UFlat/8 380267 379233 10000 1073.2MB/s txt3 BM_UFlat/9 529495 528194 7957 870.0MB/s txt4 BM_UFlat/10 31844 31784 100000 3.5GB/s pb BM_UFlat/11 146822 146476 28737 1.2GB/s gaviota BM_UFlat/12 10784 10762 392176 2.1GB/s cp BM_UFlat/13 5528 5518 760934 1.9GB/s c BM_UFlat/14 1721 1719 2449291 2.0GB/s lsp BM_UFlat/15 673304 671774 6255 1.4GB/s xls BM_UFlat/16 155 155 27092003 1.2GB/s xls_200 BM_UFlat/17 230424 229902 18285 2.1GB/s bin BM_UFlat/18 185 184 22818199 1033.9MB/s bin_200 BM_UFlat/19 21035 20996 200765 1.7GB/s sum BM_UFlat/20 2242 2238 1864380 1.8GB/s man BM_UFlatSink/0 33487 33405 100000 2.9GB/s html BM_UFlatSink/1 431108 430226 9764 1.5GB/s urls BM_UFlatSink/2 5927 5916 648112 19.4GB/s jpg BM_UFlatSink/3 123 122 34704423 1.5GB/s jpg_200 BM_UFlatSink/4 6472 6461 653462 14.8GB/s pdf BM_UFlatSink/5 164309 163988 25567 2.3GB/s html4 BM_UFlatSink/6 138274 138020 30311 1050.9MB/s txt1 BM_UFlatSink/7 120844 120637 34708 989.6MB/s txt2 BM_UFlatSink/8 371046 370366 10000 1098.9MB/s txt3 BM_UFlatSink/9 510021 508982 8269 902.9MB/s txt4 BM_UFlatSink/10 30889 30844 100000 3.6GB/s pb BM_UFlatSink/11 140752 140521 29903 1.2GB/s gaviota BM_UFlatSink/12 10162 10146 413600 2.3GB/s cp BM_UFlatSink/13 5264 5256 762398 2.0GB/s c BM_UFlatSink/14 1622 1619 2606069 2.1GB/s lsp BM_UFlatSink/15 646897 645756 6512 1.5GB/s xls BM_UFlatSink/16 150 150 28223595 1.2GB/s xls_200 BM_UFlatSink/17 226096 225650 18629 2.1GB/s bin BM_UFlatSink/18 185 184 22907935 1035.3MB/s bin_200 BM_UFlatSink/19 21369 21335 198881 1.7GB/s sum BM_UFlatSink/20 2139 2136 1953637 1.8GB/s man
Snappy, a fast compressor/decompressor. Introduction ============ Snappy is a compression/decompression library. It does not aim for maximum compression, or compatibility with any other compression library; instead, it aims for very high speeds and reasonable compression. For instance, compared to the fastest mode of zlib, Snappy is an order of magnitude faster for most inputs, but the resulting compressed files are anywhere from 20% to 100% bigger. (For more information, see "Performance", below.) Snappy has the following properties: * Fast: Compression speeds at 250 MB/sec and beyond, with no assembler code. See "Performance" below. * Stable: Over the last few years, Snappy has compressed and decompressed petabytes of data in Google's production environment. The Snappy bitstream format is stable and will not change between versions. * Robust: The Snappy decompressor is designed not to crash in the face of corrupted or malicious input. * Free and open source software: Snappy is licensed under a BSD-type license. For more information, see the included COPYING file. Snappy has previously been called "Zippy" in some Google presentations and the like. Performance =========== Snappy is intended to be fast. On a single core of a Core i7 processor in 64-bit mode, it compresses at about 250 MB/sec or more and decompresses at about 500 MB/sec or more. (These numbers are for the slowest inputs in our benchmark suite; others are much faster.) In our tests, Snappy usually is faster than algorithms in the same class (e.g. LZO, LZF, FastLZ, QuickLZ, etc.) while achieving comparable compression ratios. Typical compression ratios (based on the benchmark suite) are about 1.5-1.7x for plain text, about 2-4x for HTML, and of course 1.0x for JPEGs, PNGs and other already-compressed data. Similar numbers for zlib in its fastest mode are 2.6-2.8x, 3-7x and 1.0x, respectively. More sophisticated algorithms are capable of achieving yet higher compression rates, although usually at the expense of speed. Of course, compression ratio will vary significantly with the input. Although Snappy should be fairly portable, it is primarily optimized for 64-bit x86-compatible processors, and may run slower in other environments. In particular: - Snappy uses 64-bit operations in several places to process more data at once than would otherwise be possible. - Snappy assumes unaligned 32- and 64-bit loads and stores are cheap. On some platforms, these must be emulated with single-byte loads and stores, which is much slower. - Snappy assumes little-endian throughout, and needs to byte-swap data in several places if running on a big-endian platform. Experience has shown that even heavily tuned code can be improved. Performance optimizations, whether for 64-bit x86 or other platforms, are of course most welcome; see "Contact", below. Usage ===== Note that Snappy, both the implementation and the main interface, is written in C++. However, several third-party bindings to other languages are available; see the home page at http://google.github.io/snappy/ for more information. Also, if you want to use Snappy from C code, you can use the included C bindings in snappy-c.h. To use Snappy from your own C++ program, include the file "snappy.h" from your calling file, and link against the compiled library. There are many ways to call Snappy, but the simplest possible is snappy::Compress(input.data(), input.size(), &output); and similarly snappy::Uncompress(input.data(), input.size(), &output); where "input" and "output" are both instances of std::string. There are other interfaces that are more flexible in various ways, including support for custom (non-array) input sources. See the header file for more information. Tests and benchmarks ==================== When you compile Snappy, snappy_unittest is compiled in addition to the library itself. You do not need it to use the compressor from your own library, but it contains several useful components for Snappy development. First of all, it contains unit tests, verifying correctness on your machine in various scenarios. If you want to change or optimize Snappy, please run the tests to verify you have not broken anything. Note that if you have the Google Test library installed, unit test behavior (especially failures) will be significantly more user-friendly. You can find Google Test at http://github.com/google/googletest You probably also want the gflags library for handling of command-line flags; you can find it at http://gflags.github.io/gflags/ In addition to the unit tests, snappy contains microbenchmarks used to tune compression and decompression performance. These are automatically run before the unit tests, but you can disable them using the flag --run_microbenchmarks=false if you have gflags installed (otherwise you will need to edit the source). Finally, snappy can benchmark Snappy against a few other compression libraries (zlib, LZO, LZF, FastLZ and QuickLZ), if they were detected at configure time. To benchmark using a given file, give the compression algorithm you want to test Snappy against (e.g. --zlib) and then a list of one or more file names on the command line. The testdata/ directory contains the files used by the microbenchmark, which should provide a reasonably balanced starting point for benchmarking. (Note that baddata[1-3].snappy are not intended as benchmarks; they are used to verify correctness in the presence of corrupted data in the unit test.) Contact ======= Snappy is distributed through GitHub. For the latest version, a bug tracker, and other information, see http://google.github.io/snappy/ or the repository at https://github.com/google/snappy
Description
Languages
C++
89.5%
CMake
5.4%
C
2.6%
Starlark
2.5%