Skip to content

fix(cpp): handle TS2DIFF float prefixes in batch decode - #901

Open
kkzi wants to merge 1 commit into
apache:developfrom
kkzi:fix/cpp-ts2diff-float-double-batch-prefix
Open

fix(cpp): handle TS2DIFF float prefixes in batch decode#901
kkzi wants to merge 1 commit into
apache:developfrom
kkzi:fix/cpp-ts2diff-float-double-batch-prefix

Conversation

@kkzi

@kkzi kkzi commented Aug 7, 2026

Copy link
Copy Markdown

Summary

  • make C++ FLOAT/DOUBLE TS_2DIFF batch decoding consume the Java-compatible scale/overflow prefix
  • route batch reads through the existing segment-aware scalar decoders
  • add regression coverage for multiple FLOAT segments and a DOUBLE overflow bitmap prefix

Root cause

The FLOAT/DOUBLE batch overrides delegated directly to the INT32/INT64 TS_2DIFF batch decoders and then bit-cast the results. Those integer decoders expect a delta-block header at the current stream position, but FLOAT/DOUBLE segments place a maxPointNumber or overflow-bitmap prefix before that header. The prefix was therefore decoded as block metadata, leaving the stream and decoder state inconsistent. In tree-reader batch paths this could result in unbounded decoding at end-of-input.

The scalar read_float and read_double implementations already consume the prefix and apply scale/overflow handling. Reusing those implementations restores correctness for normal, overflow, and legacy raw segments. This intentionally gives up the integer SIMD fast path for FLOAT/DOUBLE until a prefix-aware optimized implementation is available.

Tests

  • FloatDoubleTS2DIFFCodecTest.*
  • TS2DIFFCodecTest.*
  • FloatTS2DIFFEncoderResetTest.*
  • EncodingCoverage.TS2DIFF*
  • TreeQueryByRowTest.QueryByRow_TabletMultiType_PartialPaths
  • TsFileWriterTest.WriteDiffrentTypeCombination
  • clang-format --dry-run --Werror cpp/src/encoding/ts2diff_decoder.h cpp/test/encoding/ts2diff_codec_test.cc

Fixes #900

@ColinLeeo

Copy link
Copy Markdown
Contributor

Thanks for tracking this down.

The root-cause analysis is clear, and the new implementation correctly handles Java-compatible prefixes, including overflow prefixes and reads spanning multiple segments.

I found one blocking compatibility issue, though: routing FLOAT/DOUBLE batch reads through the scalar decoder regresses legacy raw segments. The scalar prefix detector can misclassify a valid raw header, after which the decoder gets an invalid bit_width_ and spins at end-of-input.

I reproduced this for both FLOAT and DOUBLE by encoding 129 sequential raw bit patterns with IntTS2DIFFEncoder / LongTS2DIFFEncoder, then reading them in small batches through the corresponding floating-point decoder. The PR head hangs, while the parent implementation completes successfully.

Could we preserve the integer batch path for legacy raw segments, or make the prefix detection unambiguous before switching to the scalar path? It would also be good to add legacy raw batch regression tests for both types.

@ColinLeeo
ColinLeeo self-requested a review August 10, 2026 08:25

@ColinLeeo ColinLeeo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The overall fix direction looks good, but the legacy raw segment compatibility issue is not fully addressed yet.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] C++ FLOAT/DOUBLE TS_2DIFF batch decoding skips segment prefix

2 participants