apache / apache/datasketches-python

vector_of_kll_floats_sketches.get_quantiles() returns wrong values with float32

未关闭
#63 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Jupyter Notebook
星标
46
派生
9
PR 合并指标
30 天内没有已合并 PR

描述

```python
#!/usr/bin/env python3
"""
Minimal example: vector_of_kll_floats_sketches.get_quantiles() returns WRONG VALUES with float32
"""
import numpy as np
from datasketches import vector_of_kll_floats_sketches

# Create test data: 1000 samples between -100 and -10
np.random.seed(42)
test_data = np.random.uniform(-100, -10, size=(1000, 1)).astype(np.float32)

print("Test data: 1000 samples between -100 and -10")
print(f"True min: {test_data.min():.2f}, True max: {test_data.max():.2f}")

# Create sketch and add data
kll = vector_of_kll_floats_sketches(200, 1)
kll.update(test_data)

# Request p0.0001 (should be ~-100) and p0.9999 (should be ~-10)
ranks_list = [0.0001, 0.9999]
ranks_array32 = np.array(ranks_list, dtype=np.float32)
ranks_array64 = np.array(ranks_list, dtype=np.float64)

print("\n" + "="*60)
print("BUG: numpy array with dtype=np.float32 returns WRONG quantiles")
print("="*60)

quants_array = kll.get_quantiles(ranks_array32)
print(f"\nWith numpy array with dtype=np.float32: {ranks_array32}")
print(f" p0.0001 = {quants_array[0][0]:.2f} (expected: ~-100)")
print(f" p0.9999 = {quants_array[0][1]:.2f} (expected: ~-10)")
print(f" ✗ WRONG: Both values near minimum!")

quants_array64 = kll.get_quantiles(ranks_array64)
print(f"\nWith numpy array with dtype=np.float64: {ranks_array64}")
print(f" p0.0001 = {quants_array64[0][0]:.2f} (expected: ~-100)")
print(f" p0.9999 = {quants_array64[0][1]:.2f} (expected: ~-10)")
print(f" ✓ CORRECT")

```

```
Test data: 1000 samples between -100 and -10
True min: -99.58, True max: -10.03

============================================================
BUG: numpy array with dtype=np.float32 returns WRONG quantiles
============================================================

With numpy array with dtype=np.float32: [1.000e-04 9.999e-01]
p0.0001 = -98.69 (expected: ~-100)
p0.9999 = -99.50 (expected: ~-10)
✗ WRONG: Both values near minimum!

With numpy array with dtype=np.float64: [1.000e-04 9.999e-01]
p0.0001 = -99.50 (expected: ~-100)
p0.9999 = -10.28 (expected: ~-10)
✓ CORRECT
```

贡献指南

这个仓库没有索引到贡献指南

调研方向

从使用 vector_of_kll_floats_sketches.get_quantiles() 的 Python 复现开始,比较 float32 和 float64 排名数组。跟踪 float32 排名的处理,并添加一个覆盖接近 0 和 1 的排名的回归测试;完成的标准是,对于该示例,float32 和 float64 返回等价的分位数。

由索引模型根据 Issue 内容生成。

评估

技术栈
numpy, python
领域
data
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
45/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。