fce661fa8cfec792dd1ef7fee52320188349feae
This is a performance tuning change that should not affect correctness. On perflab with FDO on Haswell the performance gain is 21,776ns before vs 21,255ns after, about 2.4%. (Using geometric means.) SAMPLE PERFORMANCE with FDO on HASWELL (NEW) Benchmark Time(ns) CPU(ns) Iterations ------------------------------------------------ BM_UFlat/0 37366 37279 100000 2.6GB/s html BM_UFlat/1 471153 470204 8975 1.4GB/s urls BM_UFlat/2 6116 6105 639496 18.8GB/s jpg BM_UFlat/3 123 123 34709908 1.5GB/s jpg_200 BM_UFlat/4 6724 6714 623318 14.2GB/s pdf BM_UFlat/5 183122 182722 23138 2.1GB/s html4 BM_UFlat/6 144981 144689 29384 1002.5MB/s txt1 BM_UFlat/7 125939 125691 33423 949.8MB/s txt2 BM_UFlat/8 383101 382241 10000 1064.7MB/s txt3 BM_UFlat/9 527824 526606 7958 872.6MB/s txt4 BM_UFlat/10 34849 34790 100000 3.2GB/s pb BM_UFlat/11 150213 149937 28131 1.1GB/s gaviota BM_UFlat/12 10850 10830 393231 2.1GB/s cp BM_UFlat/13 5532 5523 735739 1.9GB/s c BM_UFlat/14 1698 1695 2478035 2.0GB/s lsp BM_UFlat/15 678396 676917 6200 1.4GB/s xls BM_UFlat/16 155 155 26909789 1.2GB/s xls_200 BM_UFlat/17 241235 240698 17416 2.0GB/s bin BM_UFlat/18 183 183 23000841 1043.5MB/s bin_200 BM_UFlat/19 21461 21424 193275 1.7GB/s sum BM_UFlat/20 2232 2228 1887191 1.8GB/s man BM_UFlatSink/0 42272 42199 98528 2.3GB/s html BM_UFlatSink/1 460814 459898 9092 1.4GB/s urls BM_UFlatSink/2 5558 5547 768629 20.7GB/s jpg BM_UFlatSink/3 124 123 33629141 1.5GB/s jpg_200 BM_UFlatSink/4 6634 6621 629989 14.4GB/s pdf BM_UFlatSink/5 182883 182491 23030 2.1GB/s html4 BM_UFlatSink/6 143269 142964 29410 1014.5MB/s txt1 BM_UFlatSink/7 127041 126809 33136 941.4MB/s txt2 BM_UFlatSink/8 384367 383577 10000 1061.0MB/s txt3 BM_UFlatSink/9 529979 528890 7898 868.9MB/s txt4 BM_UFlatSink/10 41154 41075 100000 2.7GB/s pb BM_UFlatSink/11 146446 146155 28742 1.2GB/s gaviota BM_UFlatSink/12 11939 11918 352663 1.9GB/s cp BM_UFlatSink/13 5430 5421 770451 1.9GB/s c BM_UFlatSink/14 1665 1662 2538921 2.1GB/s lsp BM_UFlatSink/15 666840 665617 6309 1.4GB/s xls BM_UFlatSink/16 152 152 27639460 1.2GB/s xls_200 BM_UFlatSink/17 240076 239573 17643 2.0GB/s bin BM_UFlatSink/18 183 182 23128210 1046.0MB/s bin_200 BM_UFlatSink/19 22570 22528 185839 1.6GB/s sum BM_UFlatSink/20 2183 2180 1899526 1.8GB/s man SAMPLE PERFORMANCE with FDO on HASWELL (OLD) Benchmark Time(ns) CPU(ns) Iterations ------------------------------------------------ BM_UFlat/0 37041 36990 100000 2.6GB/s html BM_UFlat/1 471384 470574 8930 1.4GB/s urls BM_UFlat/2 5997 5986 722354 19.2GB/s jpg BM_UFlat/3 124 123 34964717 1.5GB/s jpg_200 BM_UFlat/4 6850 6838 621414 13.9GB/s pdf BM_UFlat/5 182578 182271 23001 2.1GB/s html4 BM_UFlat/6 148338 147989 28132 980.1MB/s txt1 BM_UFlat/7 130682 130471 32347 915.0MB/s txt2 BM_UFlat/8 397420 396553 10000 1026.3MB/s txt3 BM_UFlat/9 550126 548872 7736 837.2MB/s txt4 BM_UFlat/10 35013 34958 100000 3.2GB/s pb BM_UFlat/11 152270 151889 27508 1.1GB/s gaviota BM_UFlat/12 11117 11096 379059 2.1GB/s cp BM_UFlat/13 5812 5801 725240 1.8GB/s c BM_UFlat/14 1780 1777 2383982 2.0GB/s lsp BM_UFlat/15 707871 706139 5946 1.4GB/s xls BM_UFlat/16 157 157 26889747 1.2GB/s xls_200 BM_UFlat/17 239160 238556 17512 2.0GB/s bin BM_UFlat/18 181 180 23326040 1057.5MB/s bin_200 BM_UFlat/19 22706 22656 186285 1.6GB/s sum BM_UFlat/20 2319 2315 1813186 1.7GB/s man BM_UFlatSink/0 42657 42574 99000 2.2GB/s html BM_UFlatSink/1 466316 465262 9036 1.4GB/s urls BM_UFlatSink/2 6873 6859 648525 16.7GB/s jpg BM_UFlatSink/3 124 124 34434643 1.5GB/s jpg_200 BM_UFlatSink/4 6804 6790 624282 14.0GB/s pdf BM_UFlatSink/5 185468 185062 22746 2.1GB/s html4 BM_UFlatSink/6 148511 148209 28284 978.6MB/s txt1 BM_UFlatSink/7 130865 130607 32144 914.0MB/s txt2 BM_UFlatSink/8 393931 392983 10000 1035.6MB/s txt3 BM_UFlatSink/9 545548 544275 7740 844.3MB/s txt4 BM_UFlatSink/10 41659 41584 100000 2.7GB/s pb BM_UFlatSink/11 152062 151721 27854 1.1GB/s gaviota BM_UFlatSink/12 11987 11968 350909 1.9GB/s cp BM_UFlatSink/13 5652 5641 743280 1.8GB/s c BM_UFlatSink/14 1728 1725 2446140 2.0GB/s lsp BM_UFlatSink/15 687879 686231 6138 1.4GB/s xls BM_UFlatSink/16 155 155 27254484 1.2GB/s xls_200 BM_UFlatSink/17 240689 240083 17450 2.0GB/s bin BM_UFlatSink/18 183 182 22932858 1046.8MB/s bin_200 BM_UFlatSink/19 22718 22674 185207 1.6GB/s sum BM_UFlatSink/20 2272 2268 1851664 1.7GB/s man
Snappy, a fast compressor/decompressor. Introduction ============ Snappy is a compression/decompression library. It does not aim for maximum compression, or compatibility with any other compression library; instead, it aims for very high speeds and reasonable compression. For instance, compared to the fastest mode of zlib, Snappy is an order of magnitude faster for most inputs, but the resulting compressed files are anywhere from 20% to 100% bigger. (For more information, see "Performance", below.) Snappy has the following properties: * Fast: Compression speeds at 250 MB/sec and beyond, with no assembler code. See "Performance" below. * Stable: Over the last few years, Snappy has compressed and decompressed petabytes of data in Google's production environment. The Snappy bitstream format is stable and will not change between versions. * Robust: The Snappy decompressor is designed not to crash in the face of corrupted or malicious input. * Free and open source software: Snappy is licensed under a BSD-type license. For more information, see the included COPYING file. Snappy has previously been called "Zippy" in some Google presentations and the like. Performance =========== Snappy is intended to be fast. On a single core of a Core i7 processor in 64-bit mode, it compresses at about 250 MB/sec or more and decompresses at about 500 MB/sec or more. (These numbers are for the slowest inputs in our benchmark suite; others are much faster.) In our tests, Snappy usually is faster than algorithms in the same class (e.g. LZO, LZF, FastLZ, QuickLZ, etc.) while achieving comparable compression ratios. Typical compression ratios (based on the benchmark suite) are about 1.5-1.7x for plain text, about 2-4x for HTML, and of course 1.0x for JPEGs, PNGs and other already-compressed data. Similar numbers for zlib in its fastest mode are 2.6-2.8x, 3-7x and 1.0x, respectively. More sophisticated algorithms are capable of achieving yet higher compression rates, although usually at the expense of speed. Of course, compression ratio will vary significantly with the input. Although Snappy should be fairly portable, it is primarily optimized for 64-bit x86-compatible processors, and may run slower in other environments. In particular: - Snappy uses 64-bit operations in several places to process more data at once than would otherwise be possible. - Snappy assumes unaligned 32- and 64-bit loads and stores are cheap. On some platforms, these must be emulated with single-byte loads and stores, which is much slower. - Snappy assumes little-endian throughout, and needs to byte-swap data in several places if running on a big-endian platform. Experience has shown that even heavily tuned code can be improved. Performance optimizations, whether for 64-bit x86 or other platforms, are of course most welcome; see "Contact", below. Usage ===== Note that Snappy, both the implementation and the main interface, is written in C++. However, several third-party bindings to other languages are available; see the home page at http://google.github.io/snappy/ for more information. Also, if you want to use Snappy from C code, you can use the included C bindings in snappy-c.h. To use Snappy from your own C++ program, include the file "snappy.h" from your calling file, and link against the compiled library. There are many ways to call Snappy, but the simplest possible is snappy::Compress(input.data(), input.size(), &output); and similarly snappy::Uncompress(input.data(), input.size(), &output); where "input" and "output" are both instances of std::string. There are other interfaces that are more flexible in various ways, including support for custom (non-array) input sources. See the header file for more information. Tests and benchmarks ==================== When you compile Snappy, snappy_unittest is compiled in addition to the library itself. You do not need it to use the compressor from your own library, but it contains several useful components for Snappy development. First of all, it contains unit tests, verifying correctness on your machine in various scenarios. If you want to change or optimize Snappy, please run the tests to verify you have not broken anything. Note that if you have the Google Test library installed, unit test behavior (especially failures) will be significantly more user-friendly. You can find Google Test at http://github.com/google/googletest You probably also want the gflags library for handling of command-line flags; you can find it at http://gflags.github.io/gflags/ In addition to the unit tests, snappy contains microbenchmarks used to tune compression and decompression performance. These are automatically run before the unit tests, but you can disable them using the flag --run_microbenchmarks=false if you have gflags installed (otherwise you will need to edit the source). Finally, snappy can benchmark Snappy against a few other compression libraries (zlib, LZO, LZF, FastLZ and QuickLZ), if they were detected at configure time. To benchmark using a given file, give the compression algorithm you want to test Snappy against (e.g. --zlib) and then a list of one or more file names on the command line. The testdata/ directory contains the files used by the microbenchmark, which should provide a reasonably balanced starting point for benchmarking. (Note that baddata[1-3].snappy are not intended as benchmarks; they are used to verify correctness in the presence of corrupted data in the unit test.) Contact ======= Snappy is distributed through GitHub. For the latest version, a bug tracker, and other information, see http://google.github.io/snappy/ or the repository at https://github.com/google/snappy
Description
Languages
C++
89.5%
CMake
5.4%
C
2.6%
Starlark
2.5%