← 返回日报
精读 预计 2 分钟

ALP: Adaptive Lossless Floating-Point Compression

摘要

仓库包含 ALP 算法的源代码和基准测试,用于 IEEE 754 浮点数据的无损压缩。算法通过将十进制浮点映射为整数并使用 FastLanes Frame-of-Reference 编码处理,以及对高精度浮点左部使用字典编码、右部保持未压缩。基准结果显示其在解码速度、压缩比和压缩速度上均优于其他方案。还提供复现脚本、仓库结构说明、ACM 复现报告链接,并列出在 DuckDB、FastLanes 等系统中的集成案例。

荐读理由

源码和 benchmark 完整开源,能直接拉到项目里测浮点列的压缩速度与比率,DuckDB 已集成可验证

原文

ALP: Adaptive Lossless Floating-Point Compression

Authors: Azim Afroozeh, Leonardo Kuffó, Peter Boncz Conference: ACM SIGMOD 2024


GitHub What is this repo?

This repository contains the source code and benchmarks for the paper ALP: Adaptive Lossless Floating-Point Compression, published at ACM SIGMOD 2024.

ALP is a state-of-the-art lossless compression algorithm designed for IEEE 754 floating-point data. It encodes data by exploiting two common patterns found in real-world floating-point values:

  • Decimal Floating-Point Numbers: A large portion of floats/doubles in real-world datasets are decimals. ALP maps these values into integers by multiplying the number by a power of 10 and then compressing the result using a FastLanes variant of Frame-of-Reference encoding1, which is SIMD-friendly. Example: the number 10.12 becomes 1012 and is then fed to the FastLanes encoder.

  • High-Precision Floating-Point Numbers: The remaining values are typically high-precision floats/doubles. ALP targets compression opportunities in only the left part of these values, which it compresses using FastLanes dictionary encoding. The right part is left uncompressed, as it is required to preserve high precision and is often highly random and incompressible.


📊 How does ALP perform?

ALP Results

These results highlight ALP’s superior performance across all three key metrics of a compression algorithm: Decoding Speed, Compression Ratio, and Compression Speed—outperforming other schemes in every category.


🧪 How to Reproduce Results

Just run the following script:

./publication/script/master_script.sh

For more information on reproducing our benchmarks, refer to our guide here, or read the official ACM reproducibility report: https://dl.acm.org/doi/10.1145/3687998.3717057


🏅 ACM Artifacts & Awards

We are happy to share that we participated in the SIGMOD Availability & Reproducibility Initiative, and our paper earned all three badges:

ACM Artifacts Available ACM Artifacts Evaluated ACM Results Reproduced

🎉 We're also proud to share that ALP won the SIGMOD Best Artifact Award!

Trophy


⏱️ Want to Benchmark Your Dataset?

Check out our guide: How to Benchmark Your Dataset It explains how to run ALP on your own data.


🗂️ Repository Structure

  • src/: Core implementation of ALP and ALP_RD

  • benchmarks/: Benchmarking tools and datasets

  • include/: Header files for integration

  • scripts/: Utility scripts for data processing

  • test/: Unit tests

  • publication/: Publications and supplementary materials


📚 Publications


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.


📬 Contact

If you have questions, want to contribute, or just want to stay up to date with ALP and related projects, join our community on Discord: Join us on Discord Community Status


🧩 Used By

ALP has been integrated into the following systems:


Footnotes

  1. Learn more about FastLanes here: https://github.com/cwida/fastlanes
Lobsters · 1 赞 · 0 评 讨论 → 阅读原文 →

这条对你有帮助吗?