inserting data slow when insert data in bulk
- Lenguaje dominante
- Java
- Estrellas
- 6.4k
- Forks
- 1.2k
- Merge medio
- 1 d 23 h
- PR fusionados (30 d)
- 115
Descripción
**environment:**
version: 1.0.0
deployment: all in one, one confignode, one datanode
host: cpu: 16core, mem:16G, disk: ssd, os: ubuntu 22.04
client sdk: python, use normal tablet
**background:**
the measurements number of a batch of data is huge, in my usage scenario, sometimes it is 200,000.
**experiment:**
I tested the insertion performance.
Insertion takes 2.5s-4s when all measurement points are inserted in one database and the measurements number of a batch is 200,000.
Insertion takes 0.3s when all measurement points are inserted in 10 databases of one instance and 10 threads, one thread insert 20,000 measurements once.
I looked at the usage performance of host, cpu only used under 200%, memory is still remain, disk io is 8MB/s.
None of this is a bottleneck. so i strace the process:
```
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ------------------
79.04 52.355396 1545 33867 4478 futex
8.88 5.880482 2816 2088 epoll_pwait2
5.25 3.480728 102374 34 15 restart_syscall
4.60 3.049973 7682 397 epoll_wait
1.84 1.217815 1363 893 recvfrom
0.19 0.124981 17 7105 writev
0.10 0.065069 10 6045 read
0.03 0.023011 442 52 pwrite64
0.02 0.013909 5 2750 write
0.01 0.005350 137 39 fdatasync
```
in strace log, many lock timeout:
```
335339 16:00:27.993273 futex(0x7f1b240f6028, FUTEX_WAKE_PRIVATE, 1) = 0 <0.000011>
335339 16:00:27.993317 futex(0x7f1b240f6078, FUTEX_WAIT_BITSET_PRIVATE, 0, {tv_sec=359621, tv_nsec=748782361}, FUTEX_BITSET_MATCH_ANY
335381 16:00:27.995534 <... futex resumed>) = -1 ETIMEDOUT (Connection timed out) <0.200098>
335381 16:00:27.995648 futex(0x7f19e0025718, FUTEX_WAKE_PRIVATE, 1) = 0 <0.000100>
335381 16:00:27.995913 futex(0x7f19e0025768, FUTEX_WAIT_BITSET_PRIVATE, 0, {tv_sec=359621, tv_nsec=901319853}, FUTEX_BITSET_MATCH_ANY
335454 16:00:27.996630 <... futex resumed>) = -1 ETIMEDOUT (Connection timed out) <0.100112>
335454 16:00:27.996656 futex(0x7f19e00b70d8, FUTEX_WAKE_PRIVATE, 1) = 0 <0.000009>
335454 16:00:27.996762 futex(0x7f19e00b7128, FUTEX_WAIT_BITSET_PRIVATE, 0, {tv_sec=359621, tv_nsec=802178374}, FUTEX_BITSET_MATCH_ANY
```
so i traceing java process use jstack, but I can't read the contents, but there is code related to locks.
**question:**
1. In cases where there are so many measurement points, How do I need to optimize it? I look at the documentation and there are too many configuration parameters. https://iotdb.apache.org/UserGuide/Master/Reference/Common-Config-Manual.html
2. Can it improve performance, if i use cluster mode.
Guía de contribución
Línea de trabajo
El informe no menciona archivos ni pruebas del repositorio; comienza reproduciendo la inserción masiva con el despliegue de IoTDB indicado, el cliente de Python, la salida de strace y la captura de jstack. Lee el Common Config Manual enlazado y compara ejecuciones con una sola base de datos y en modo clúster. Se considerará completado cuando se identifique el bloqueo o la configuración limitante y se documente una optimización verificada o un resultado de rendimiento del clúster.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- java, python
- Área
- databases, performance
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 25/100