git.ipfire.org Git - thirdparty/zstd.git/log

]> git.ipfire.org Git - thirdparty/zstd.git/log

projects / thirdparty / zstd.git / log

summary | shortlog | log | commit | commitdiff | tree
first ⋅ prev ⋅ next

commit | commitdiff | tree

Yann Collet [Wed, 3 Aug 2022 19:39:35 +0000 (21:39 +0200)]

fileio_types.h : avoid dependency on mem.h

fileio_types.h cannot be parsed by itself
because it relies on basic types defined in `lib/common/mem.h`.
As for #3231, it likely wasn't detected because `mem.h` was probably included before within target files.
But this is not proper.

A "easy" solution would be to add the missing include,
but each dependency should be considered "bad" by default,
and only allowed if it brings some tangible value.

In this case, since these types are only used to declare internal structure variables
which are effectively only flags,
I believe it's really not valuable to add a dependency on `mem.h` for this purpose
while the standard `int` type can do the same job.

I was expecting some compiler warnings following this change,
but it turns out we don't use `-Wconversion` by default on `zstd` source code,
so there is none.

Nevertheless, I enabled `-Wconversion` locally and proceeded to fix a few conversion warnings in the process.

Adding `-Wconversion` to the list of flags used for `zstd` is something I would be favorable over the long term,
but it cannot be done overnight,
because the nb of places where this warning is triggered is daunting.
Better progressively reduce the nb of triggered `-Wconversion` warnings before enabling this flag by default.

commit | commitdiff | tree

Felix Handte [Tue, 2 Aug 2022 16:34:04 +0000 (12:34 -0400)]

Merge pull request #3196 from mileshu/dev

[T124890272] Mark 2 Obsolete Functions(ZSTD_copy*Ctx) Deprecated in Zstd

commit | commitdiff | tree

Miles Hu [Tue, 2 Aug 2022 05:52:47 +0000 (22:52 -0700)]

Merge branch 'dev' of https://github.com/mileshu/zstd into dev

commit | commitdiff | tree

Miles HU [Wed, 13 Jul 2022 18:00:05 +0000 (11:00 -0700)]

[T124890272] Mark 2 Obsolete Functions(ZSTD_copy*Ctx) Deprecated in Zstd

The discussion for this task is here: facebook/zstd#3128.

This task can probably be scoped to the first part: marking these functions deprecated.
We'll later look at removal when we roll out v1.6.0.

commit | commitdiff | tree

Nick Terrell [Mon, 1 Aug 2022 18:52:14 +0000 (11:52 -0700)]

Deprecate ZSTD_getDecompressedSize() (#3225)

Fixes #3158.

Mark ZSTD_getDecompressedSize() as deprecated and replaced by ZSTD_getFrameContentSize().

commit | commitdiff | tree

Elliot Gorokhovsky [Mon, 1 Aug 2022 18:04:57 +0000 (14:04 -0400)]

Merge pull request #3220 from embg/issue3200

Disallow empty string as argument for --output-dir-flat and --output-dir-mirror

commit | commitdiff | tree

Qiongsi Wu [Mon, 1 Aug 2022 17:41:24 +0000 (13:41 -0400)]

Fix hash4Ptr for big endian (#3227)

commit | commitdiff | tree

Yonatan Komornik [Fri, 29 Jul 2022 23:13:07 +0000 (16:13 -0700)]

stdin multiple file fixes (#3222)

* Fixes for https://github.com/facebook/zstd/issues/3206 - bugs when handling stdin as part of multiple files.

* new line at end of multiple-files.sh

commit | commitdiff | tree

Elliot Gorokhovsky [Fri, 29 Jul 2022 21:44:22 +0000 (14:44 -0700)]

Disallow empty output directory

commit | commitdiff | tree

Tom Wang [Fri, 29 Jul 2022 19:51:58 +0000 (12:51 -0700)]

Add warning when multi-thread decompression is requested (#3208)

When user pass in argument for both decompression and multi-thread, print a warning message
to indicate that multi-threaded decompression is not supported.

* Add warning when multi-thread decompression is requested
* add test case for multi-threaded decoding warning
Expectation is for -d -T0 we will not throw any warning,
and see warning for any other -d -T(>1) inputs

commit | commitdiff | tree

Chris Burgess [Fri, 29 Jul 2022 19:22:46 +0000 (15:22 -0400)]

Fix small file passthrough (#3215)

commit | commitdiff | tree

orbea [Fri, 29 Jul 2022 19:22:10 +0000 (12:22 -0700)]

zlibWrapper: Update for zlib 1.2.12 (#3217)

In zlib 1.2.12 the OF macro was changed to _Z_OF breaking any
project that used zlibWrapper. To fix this the OF has been
changed to _Z_OF everywhere and _Z_OF is defined as OF in the
case it is not yet defined for zlib 1.2.11 and older.

Fixes: https://github.com/facebook/zstd/issues/3216

commit | commitdiff | tree

Qiongsi Wu [Fri, 29 Jul 2022 19:21:59 +0000 (15:21 -0400)]

[AIX] Fix Compiler Flags and Bugs on AIX to Pass All Tests (#3219)

* Fixing compiler warnings

* Replace the old -s flag with the -Wl,-s flag

* Fixing compiler warnings

* Fixing the linker strip flag and tests/code not working as expected on AIX

commit | commitdiff | tree

Elliot Gorokhovsky [Fri, 29 Jul 2022 18:10:47 +0000 (11:10 -0700)]

Fix buffer underflow for null dir1

commit | commitdiff | tree

Jun He [Fri, 29 Jul 2022 17:28:04 +0000 (01:28 +0800)]

lib: add hint to generate more pipeline friendly code (#3138)

With statistic data of test data files of silesia
the chance of position beyond highThreshold is very
low (~1.3%@L8 in most cases, all <2.5%), and is in
"lowprob area". Add the branch hint so compiler can
get better pipiline codegen.
With this change it is observed ~1% of mozilla and
xml, and slight (0.3%~0.8%) but consistent uplift on
other files on Arm N1.

Signed-off-by: Jun He <jun.he@arm.com>
Change-Id: Id9ba1d5c767e975290b5c1bf0ecce906544f4ade

commit | commitdiff | tree

Jun He [Fri, 29 Jul 2022 17:27:20 +0000 (01:27 +0800)]

decomp: add prefetch for matched seq on aarch64 (#3164)

match is used for following sequence copy. It is
only updated when extDict is needed, which is a
low probability case. So it can be prefetched to
reduce cache miss.
The benchmarks on various Arm platforms showed
uplift from 1% ~ 14% with gcc-11/clang-14.

Signed-off-by: Jun He <jun.he@arm.com>
Change-Id: If201af4799d2455d74c79f8387404439d7f684ae

commit | commitdiff | tree

Mathew R Gordon [Fri, 29 Jul 2022 17:17:31 +0000 (11:17 -0600)]

Add transparency and optimize logo (#3218)

Make the front page look better in dark GH themes

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 26 Jul 2022 17:26:15 +0000 (13:26 -0400)]

Merge pull request #3197 from embg/docstring_clarify

Clarify benchmark chunking docstring

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 26 Jul 2022 17:19:38 +0000 (13:19 -0400)]

Merge pull request #3209 from zhuhan0/dev

[largeNbDicts] Second try at fixing decompression segfault to always create compressInstructions

commit | commitdiff | tree

Han Zhu [Wed, 20 Jul 2022 23:01:32 +0000 (16:01 -0700)]

[largeNbDicts] Second try at fixing decompression segfault to always create compressInstructions

Summary:
Freeing an uninitialized pointer is undefined behavior. This caused a segfault
when compiling the benchmark with Clang -O3 and benching decompression.

V2: always create compressInstructions but check if cctxParams is NULL before
setting CCtx params to avoid segfault.

Test Plan:
make and run

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 20 Jul 2022 20:07:04 +0000 (16:07 -0400)]

Merge pull request #3205 from zhuhan0/dev

[contrib][largeNbDicts] Fix decompression segfault; Add additional benchmark metrics

commit | commitdiff | tree

Han Zhu [Wed, 20 Jul 2022 18:14:51 +0000 (11:14 -0700)]

[largeNbDicts] Add an option to print out median speed

Summary:
Added an option -p# where -p0 (default) sets the aggregation method to fastest
speed while -p1 sets the aggregation method to median. Also added a new column
in the csv file to report this option's value.

Test Plan:
``
$ ./largeNbDicts -1 --nbDicts=1 -D ~/benchmarks/html/html_8_16K.32K.dict
~/benchmarks/html/html_8_16K/*
loading 7450 files...
created src buffer of size 83.4 MB
split input into 7450 blocks
loading dictionary /home/zhuhan/benchmarks/html/html_8_16K.32K.dict
compressing at level 1 without dictionary : Ratio=3.03  (28827863 bytes)
compressed using a 32768 bytes dictionary : Ratio=4.28  (20410262 bytes)
generating 1 dictionaries, using 0.1 MB of memory
Compression Speed : 306.0 MB/s
Fastest Speed : 310.6 MB/s

$ ./largeNbDicts -1 --nbDicts=1 -p1 -D ~/benchmarks/html/html_8_16K.32K.dict
~/benchmarks/html/html_8_16K/*
loading 7450 files...
created src buffer of size 83.4 MB
split input into 7450 blocks
loading dictionary /home/zhuhan/benchmarks/html/html_8_16K.32K.dict
compressing at level 1 without dictionary : Ratio=3.03  (28827863 bytes)
compressed using a 32768 bytes dictionary : Ratio=4.28  (20410262 bytes)
generating 1 dictionaries, using 0.1 MB of memory
Compression Speed : 306.9 MB/s
Median Speed : 298.4 MB/s
```

commit | commitdiff | tree

Han Zhu [Tue, 19 Jul 2022 23:50:28 +0000 (16:50 -0700)]

[largeNbDicts] Print more metrics into csv file

Summary:
Add column headers and data for whether it's a compression or a decompression
run, compression level, nbDicts and dictAttachPref in additional to
compr/decompr speed.

Test Plan:
Example output:

```
./largeNbDicts
Compression/Decompression,Level,nbDicts,dictAttachPref,Speed
Compression,1,1,0,300.9
Compression,1,1,1,296.4
Compression,1,1,2,307.8
Compression,1,10,0,292.3
Compression,1,100,0,293.3
Compression,3,110,0,106.0
Decompression,-1,110,-1,155.6
Decompression,-1,110,-1,709.4
Decompression,-1,120,-1,709.1
Decompression,-1,120,-1,734.6
```

commit | commitdiff | tree

Han Zhu [Tue, 19 Jul 2022 20:55:48 +0000 (13:55 -0700)]

[largeNbDicts] Fix decompression segfault in createCompressInstructions

Benchmarking decompression results in a segfault in `createCompressInstructions`
because `cctxParams` is NULL. Skip running that function if we are not benching
compression.

commit | commitdiff | tree

udayanbapat [Thu, 14 Jul 2022 18:54:34 +0000 (11:54 -0700)]

Intial commit to address 3090. Added support to decompress empty block. (#3118)

* Intial commit to address 3090. Added support to decompress empty block

* Update zstd_decompress_block.c

Addressed review comments for the case of 'set_basic'

* Update lib/decompress/zstd_decompress_block.c

Co-authored-by: Nick Terrell <nickrterrell@gmail.com>
* Update lib/decompress/zstd_decompress_block.c

Co-authored-by: Nick Terrell <nickrterrell@gmail.com>
Co-authored-by: Nick Terrell <nickrterrell@gmail.com>

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 13 Jul 2022 20:54:29 +0000 (16:54 -0400)]

Clarify -B docstring

commit | commitdiff | tree

Miles HU [Wed, 13 Jul 2022 18:00:05 +0000 (11:00 -0700)]

[T124890272] Mark 2 Obsolete Functions(ZSTD_copy*Ctx) Deprecated in Zstd

The discussion for this task is here: facebook/zstd#3128.

This task can probably be scoped to the first part: marking these functions deprecated.
We'll later look at removal when we roll out v1.6.0.

commit | commitdiff | tree

Miles HU [Tue, 12 Jul 2022 18:17:25 +0000 (11:17 -0700)]

Revert "T119975957"

This reverts commit 962746edffa5340315136af34ac3331eba82c3c8.

commit | commitdiff | tree

Miles HU [Fri, 8 Jul 2022 22:01:36 +0000 (15:01 -0700)]

T119975957

Signed-off-by: Miles HU <yuanpu@fb.com>

commit | commitdiff | tree

Felix Handte [Fri, 8 Jul 2022 20:04:39 +0000 (16:04 -0400)]

Merge pull request #3184 from htnhan/features/list_verbose_to_show_dictionary_id

zstd -lv <file> to show dictID

commit | commitdiff | tree

htnhan [Fri, 8 Jul 2022 17:20:50 +0000 (12:20 -0500)]

Detect multiple dictIDs in one file

commit | commitdiff | tree

htnhan [Wed, 6 Jul 2022 02:28:33 +0000 (21:28 -0500)]

zstd -lv <file> to show dictID

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 5 Jul 2022 17:13:34 +0000 (13:13 -0400)]

Merge pull request #3180 from nocnokneo/MSVCBuildTests

Fix ZSTD_BUILD_TESTS=ON with MSVC

commit | commitdiff | tree

Taylor Braun-Jones [Thu, 30 Jun 2022 17:20:42 +0000 (13:20 -0400)]

Fix ZSTD_BUILD_TESTS=ON build with MSVC

Fixes:

Command line error D8021 : invalid numeric argument '/Wno-deprecated-declarations'

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 29 Jun 2022 20:03:52 +0000 (13:03 -0700)]

Merge pull request #3179 from embg/1.5.3_bump

Prepare v1.5.3

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 29 Jun 2022 18:55:14 +0000 (14:55 -0400)]

make -C programs zstd.1

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 29 Jun 2022 17:11:13 +0000 (13:11 -0400)]

1.5.3 version bump

commit | commitdiff | tree

Elliot Gorokhovsky [Fri, 24 Jun 2022 15:24:43 +0000 (08:24 -0700)]

Merge pull request #3177 from embg/dms_prefetch2

Add prefetchCDictTables CCtxParam (+10-20% cold dict compression speed)

commit | commitdiff | tree

Elliot Gorokhovsky [Thu, 23 Jun 2022 20:58:03 +0000 (16:58 -0400)]

Nits

commit | commitdiff | tree

Elliot Gorokhovsky [Thu, 23 Jun 2022 01:02:07 +0000 (18:02 -0700)]

Update README.md for fuzzers (#3174)

* Update README.md for fuzzers

* Add ls corpora/*crash command

* nit

* Clarify wording and add Nick's command

* Minor clarification

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 22 Jun 2022 21:05:23 +0000 (17:05 -0400)]

Add tests

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 22 Jun 2022 15:59:28 +0000 (08:59 -0700)]

add prefetchCDictTables to largeNbDicts

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 21 Jun 2022 22:06:48 +0000 (18:06 -0400)]

Add docs

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 21 Jun 2022 15:59:27 +0000 (11:59 -0400)]

Add prefetchCDictTables CCtxParam

commit | commitdiff | tree

Yann Collet [Wed, 22 Jun 2022 18:21:09 +0000 (11:21 -0700)]

Merge pull request #3175 from facebook/fix3169

Streaming decompression can detect incorrect header ID sooner

commit | commitdiff | tree

Yann Collet [Wed, 22 Jun 2022 01:14:11 +0000 (18:14 -0700)]

Streaming decompression can detect incorrect header ID sooner

Streaming decompression used to wait for a minimum of 5 bytes before attempting decoding.
This meant that, in the case that only a few bytes (<5) were provided,
and assuming these bytes are incorrect,
there would be no error reported.
The streaming API would simply request more data, waiting for at least 5 bytes.

This PR makes it possible to detect incorrect Frame IDs as soon as the first byte is provided.

Fix #3169

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 21 Jun 2022 21:27:19 +0000 (14:27 -0700)]

"Short cache" optimization for level 1-4 DMS (+5-30% compression speed) (#3152)

* first attempt at fast DMS short cache

* significant wins for some scenarios

* fix all clang regressions

* nits

* fix 1.5% gcc11 regression on hot 110Kdict scenario

* fix CI

* nit

* Add tags to doublefast hash table

* use tags in doublefast DMS

* Fix CI

* Clean up some hardcoded logic / constants

* Switch forCCtx to an enum

* nit

* add short cache to ip+1 long search

* Move tag size into hashLog

* Minor nits

* Truncate dictionaries greater than 16MB in short cache mode

* Helper function for tag comparison

* Cap short cache hashLog at 24 to prevent overflow

* size_t dictTagsMatch -> int dictTagsMatch

* nit

* Clean up and comment dictionary truncation

* Move ZSTD_tableFillPurpose_e next to ZSTD_dictTableLoadMethod_e

* Comment and expand helper functions

* Asserts and documentation

* nit

commit | commitdiff | tree

Yann Collet [Tue, 21 Jun 2022 17:17:36 +0000 (10:17 -0700)]

Merge pull request #3170 from facebook/mesongnu99

removed gnu99 statement from meson recipe

commit | commitdiff | tree

Yann Collet [Mon, 20 Jun 2022 22:02:41 +0000 (15:02 -0700)]

removed gnu99 statement from meson recipe

commit | commitdiff | tree

Yann Collet [Sun, 19 Jun 2022 23:49:21 +0000 (16:49 -0700)]

Merge pull request #3167 from facebook/cmake_std

remove explicit standard setting from cmake script

commit | commitdiff | tree

Yann Collet [Sun, 19 Jun 2022 21:52:32 +0000 (14:52 -0700)]

removed explicit compilation standard from cmake script

it's not expected to be useful
and can actually lead to subtle side effects
such as #3163.

commit | commitdiff | tree

Yann Collet [Sun, 19 Jun 2022 21:45:49 +0000 (14:45 -0700)]

Merge pull request #3166 from facebook/warning_clockt

display a warning message when using C90 clock_t

commit | commitdiff | tree

Yann Collet [Sun, 19 Jun 2022 18:38:06 +0000 (11:38 -0700)]

display a warning message when using C90 clock_t for MT speed measurements.

commit | commitdiff | tree

Yann Collet [Sun, 19 Jun 2022 18:12:16 +0000 (11:12 -0700)]

updated documentation regarding build systems

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 15 Jun 2022 14:39:50 +0000 (07:39 -0700)]

Merge pull request #3161 from embg/largeNbDictsImprovements

[contrib] largeNbDicts bugfix + improvements

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 14 Jun 2022 23:18:49 +0000 (19:18 -0400)]

fix typo

Co-authored-by: Nick Terrell <nickrterrell@gmail.com>

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 14 Jun 2022 21:57:54 +0000 (14:57 -0700)]

Fix FILE handle leak

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 14 Jun 2022 21:52:51 +0000 (14:52 -0700)]

Support advanced API so forceCopy/forceAttach works properly

commit | commitdiff | tree

Elliot Gorokhovsky [Tue, 14 Jun 2022 00:23:33 +0000 (17:23 -0700)]

largeNbDicts bugfix + improvements

commit | commitdiff | tree

Elliot Gorokhovsky [Mon, 13 Jun 2022 18:01:43 +0000 (14:01 -0400)]

Merge pull request #3160 from danlark1/patch-1

Fix big endian ARM NEON path

commit | commitdiff | tree

Daniel Kutenin [Mon, 13 Jun 2022 08:16:24 +0000 (09:16 +0100)]

Fix big endian ARM NEON path

It is not using the NEON acceleration but the bit grouping was applied

commit | commitdiff | tree

Nick Terrell [Thu, 9 Jun 2022 20:40:51 +0000 (13:40 -0700)]

Merge pull request #3141 from JunHe77/seqDec

dec: adjust seqSymbol load on aarch64

commit | commitdiff | tree

Nick Terrell [Thu, 9 Jun 2022 20:38:30 +0000 (13:38 -0700)]

Merge pull request #3145 from JunHe77/wildcopy

common: apply two stage copy to aarch64

commit | commitdiff | tree

Elliot Gorokhovsky [Thu, 9 Jun 2022 19:35:29 +0000 (15:35 -0400)]

Merge pull request #3157 from embg/huge_dict_bugfix

Bugfix for huge dictionaries

commit | commitdiff | tree

Elliot Gorokhovsky [Thu, 9 Jun 2022 15:39:30 +0000 (11:39 -0400)]

Bugfix for huge dictionaries

commit | commitdiff | tree

Yann Collet [Wed, 8 Jun 2022 00:44:20 +0000 (17:44 -0700)]

updated --single-thread man

commit | commitdiff | tree

Nick Terrell [Mon, 6 Jun 2022 23:07:20 +0000 (16:07 -0700)]

Merge pull request #3154 from terrelln/rsyncable-speed-fix

Remove expensive assert in --rsyncable hot loop

commit | commitdiff | tree

Nick Terrell [Mon, 6 Jun 2022 18:56:13 +0000 (11:56 -0700)]

Remove expensive assert in --rsyncable hot loop

This assert slows the loop down by 10x. We can get similar
coverage by asserting at the beginning & end of the loop.

We need this fix because Debian compiles zstd with asserts
enabled. Separately, we should ask them why, and if they would
consider disabling asserts in their builds. Since we don't
optimize for assert enabled builds.

Fixes Issue #3150.

commit | commitdiff | tree

Nick Terrell [Thu, 2 Jun 2022 17:04:55 +0000 (10:04 -0700)]

Merge pull request #3147 from animalize/dev

fix leaking thread handles on Windows

commit | commitdiff | tree

Yann Collet [Thu, 2 Jun 2022 16:58:45 +0000 (09:58 -0700)]

Merge pull request #3148 from ihsinme/patch-1

simple fix

commit | commitdiff | tree

Jun He [Mon, 23 May 2022 06:25:10 +0000 (14:25 +0800)]

dec: adjust seqSymbol load on aarch64

ZSTD_seqSymbol is a structure with total of 64 bits
wide. So it can be loaded in one operation and
extract its fields by simply shifting or extracting
on aarch64.
GCC doesn't recognize this and generates more
unnecessary ldr/ldrb/ldrh operations that cause
performance drop.
With this change it is observed 2~4% uplift of
silesia and 2.5~6% of cantrbry @L8 on Arm N1.

Signed-off-by: Jun He <jun.he@arm.com>
Change-Id: I7748909204cf78a17eb9d4f2333692d53239daa8

commit | commitdiff | tree

ihsinme [Mon, 30 May 2022 11:08:19 +0000 (14:08 +0300)]

Update zstd_compress.c

commit | commitdiff | tree

Ma Lin [Mon, 30 May 2022 00:18:54 +0000 (08:18 +0800)]

fix leaking thread handles on Windows

On Windows, thread handle should be closed explicitly.

Co-authored-by: luben karavelov <luben@users.noreply.github.com>

commit | commitdiff | tree

Jun He [Wed, 25 May 2022 14:26:41 +0000 (22:26 +0800)]

common: apply two stage copy to aarch64

On aarch64 ZSTD_wildcopy uses a simple loop to do
16B based memory copy. There is existing optimized
two stage copy that can achieve better performance.
By applying this to aarch64 it is also observed ~1%
uplift in silesia corpus.

Signed-off-by: Jun He <jun.he@arm.com>
Change-Id: Ic1253308e7a8a7df2d08963ba544e086c81ce8be

commit | commitdiff | tree

Yann Collet [Tue, 24 May 2022 17:19:14 +0000 (10:19 -0700)]

Merge pull request #3143 from facebook/fixdoc_3142

fix small error in format documentation example

commit | commitdiff | tree

Nick Terrell [Tue, 24 May 2022 15:10:26 +0000 (11:10 -0400)]

Merge pull request #3139 from danlark1/dev

[lazy] Optimize ZSTD_row_getMatchMask for levels 8-10 for ARM

commit | commitdiff | tree

Yann Collet [Tue, 24 May 2022 11:47:49 +0000 (04:47 -0700)]

fix small error in format documentation example

reported by @dkcasset
fix #3142

commit | commitdiff | tree

Danila Kutenin [Mon, 23 May 2022 14:51:47 +0000 (14:51 +0000)]

Again unused error warning. Fixed

commit | commitdiff | tree

Danila Kutenin [Mon, 23 May 2022 14:49:35 +0000 (14:49 +0000)]

Move NEON version to a separate function and fix indentation

commit | commitdiff | tree

Danila Kutenin [Sun, 22 May 2022 10:50:33 +0000 (10:50 +0000)]

Disable unused variable warning

commit | commitdiff | tree

Danila Kutenin [Sun, 22 May 2022 10:34:33 +0000 (10:34 +0000)]

[lazy] Optimize ZSTD_row_getMatchMask for level 8-10

We found that movemask is not used properly or consumes too much CPU.
This effort helps to optimize the movemask emulation on ARM.

For level 8-9 we saw 3-5% improvements. For level 10 we say 1.5%
improvement.

The key idea is not to use pure movemasks but to have groups of bits.
For rowEntries == 16, 32 we are going to have groups of size 4 and 2
respectively. It means that each bit will be duplicated within the group

Then we do AND to have only one bit set in the group so that iteration
with lowering bit `a &= (a - 1)` works as well.

Also, aarch64 does not have rotate instructions for 16 bit, only for 32
and 64, that's why we see more improvements for level 8-9.

vshrn_n_u16 instruction is used to achieve that: vshrn_n_u16 shifts by
4 every u16 and narrows to 8 lower bits. See the picture below. It's
also used in
[Folly](https://github.com/facebook/folly/blob/c5702590080aa5d0e8d666d91861d64634065132/folly/container/detail/F14Table.h#L446).
It also uses 2 cycles according to Neoverse-N{1,2} guidelines.

64 bit movemask is already well optimized. We have ongoing experiments
but were not able to validate other implementations work reliably faster.

commit | commitdiff | tree

Yann Collet [Fri, 20 May 2022 17:05:16 +0000 (10:05 -0700)]

Merge pull request #3135 from averred/dev

Typo in man

commit | commitdiff | tree

Talha Khan [Fri, 20 May 2022 08:53:48 +0000 (16:53 +0800)]

Typo in man

commit | commitdiff | tree

Elliot Gorokhovsky [Thu, 12 May 2022 17:50:15 +0000 (13:50 -0400)]

Merge pull request #3127 from embg/repcode_history

Correct and clarify repcode offset history logic

commit | commitdiff | tree

Elliot Gorokhovsky [Thu, 12 May 2022 16:53:15 +0000 (12:53 -0400)]

Nits

commit | commitdiff | tree

Felix Handte [Wed, 11 May 2022 21:04:02 +0000 (17:04 -0400)]

Merge pull request #3129 from felixhandte/zstd-fast-nodict-unconditional-ip1-table-write

ZSTD_fast_noDict: Avoid Safety Check When Writing `ip1` into Table

commit | commitdiff | tree

W. Felix Handte [Wed, 11 May 2022 17:27:35 +0000 (10:27 -0700)]

Update results.csv

commit | commitdiff | tree

W. Felix Handte [Wed, 11 May 2022 16:38:20 +0000 (12:38 -0400)]

Fix Comments Slightly

commit | commitdiff | tree

W. Felix Handte [Wed, 11 May 2022 15:27:34 +0000 (11:27 -0400)]

Hoist Hash Table Writes Up into Each Match Found Block

Refactoring this way avoids the bad write in the case that `step > 4`, and
is a bit more straightforward. It also seems to perform better!

commit | commitdiff | tree

W. Felix Handte [Tue, 10 May 2022 21:29:39 +0000 (14:29 -0700)]

ZSTD_fast_noDict: Minimize Checks When Writing Hash Table for ip1

This commit avoids checking whether a hashtable write is safe in two of the
three match-found paths in `ZSTD_compressBlock_fast_noDict_generic`. This pro-
duces a ~0.5% speed-up in compression.

A comment in the code describes why we can skip this check in the other two
paths (the repcode check and the first match check in the unrolled loop).

A downside is that in the new position where we make this check, we have not
yet computed `mLength`. We therefore have to avoid writing *possibly* dangerous
positions, rather than the old check which only avoids writing *actually*
dangerous positions. This leads to a miniscule loss in ratio (remember that
this scenario can only been triggered in very negative levels or under incomp-
ressibility acceleration).

commit | commitdiff | tree

Elliot Gorokhovsky [Mon, 9 May 2022 22:26:10 +0000 (18:26 -0400)]

Nits

commit | commitdiff | tree

Elliot Gorokhovsky [Mon, 9 May 2022 21:17:11 +0000 (17:17 -0400)]

Correct and clarify repcode offset history logic

commit | commitdiff | tree

Elliot Gorokhovsky [Mon, 9 May 2022 23:48:13 +0000 (19:48 -0400)]

Merge pull request #3126 from embg/fix_freebsd_ci

Unbreak FreeBSD CI

commit | commitdiff | tree

Elliot Gorokhovsky [Mon, 9 May 2022 22:28:03 +0000 (18:28 -0400)]

Unbreak FreeBSD CI

commit | commitdiff | tree

Elliot Gorokhovsky [Thu, 5 May 2022 19:06:47 +0000 (15:06 -0400)]

Merge pull request #3114 from embg/fast_extdict_pipeline2

Software pipeline for ZSTD_compressBlock_fast_extDict

commit | commitdiff | tree

Elliot Gorokhovsky [Wed, 4 May 2022 20:05:37 +0000 (16:05 -0400)]

Update results.csv

commit | commitdiff | tree

Yann Collet [Mon, 2 May 2022 17:56:37 +0000 (10:56 -0700)]

Merge pull request #3122 from eli-schwartz/betterlinkage

meson: for internal linkage, link to both libzstd and a static copy of it

commit | commitdiff | tree

Eli Schwartz [Thu, 28 Apr 2022 22:22:55 +0000 (18:22 -0400)]

meson: for internal linkage, link to both libzstd and a static copy of it

Partial, Meson-only implementation of #2976 for non-MSVC builds.

Due to the prevalence of private symbol reuse, linking to a shared
library is simply utterly unreliable, but we still want to defer to the
shared library for installable applications. By linking to both, we can
share symbols where possible, and statically link where needed.

This means we no longer need to manually track every file that needs to
be extracted and reused.

The flip side is that MSVC completely does not support this, so for MSVC
builds we just link to a full static copy even where
-Ddefault_library=shared.

As a side benefit, by using library inclusion rather than including
extra explicit object files, the zstd program shrinks in size slightly
(~4kb).

commit | commitdiff | tree

Eli Schwartz [Tue, 10 Aug 2021 02:53:15 +0000 (22:53 -0400)]

meson: avoid rebuilding some libzstd sources in the programs

These need to be explicitly included as we use their private symbols,
but we don't need to recompile them when we can reuse the existing
objects.

Minus 7 compile steps.

commit | commitdiff | tree

Eli Schwartz [Tue, 10 Aug 2021 03:19:52 +0000 (23:19 -0400)]

meson: avoid rebuilding some libzstd files in the test programs

The poolTests program already linked to libzstd, and later to
libtestcommon with included libzstd objects. So this was redundant.

Minus 4 compile steps.

Mirror of https://github.com/facebook/zstd.git