tslearn-team / tslearn-team/tslearn

inconsistent / ambiguous values for silhouette_score for identical configuration of TimeSeriesKMeans

Open
#278 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
3.2k
Forks
384
Avg merge
3d 12h
Merged PRs (30d)
11

Description

Describe the bug
When clustering with TimeSeriesKMeans, silhouette_score yields different results even though the configuration (except for the random state obviously) is identical.

Problem is, sometimes the silhouette score for n=2 is higher than the score for n=3 and sometimes the other way around. So it is not possible to use the silhouette score for determination of optimal amount of clusters.

Is this expected behavior?

To Reproduce

from tslearn.datasets import UCR_UEA_datasets
from tslearn.clustering import TimeSeriesKMeans, silhouette_score

X_train, y_train, X_test, y_test = ds.load_dataset("CBF")

Repeat the same clustering process (n=2 clusters) for 10 times and print silhouette score:

for i in range(2,11):
    km = TimeSeriesKMeans(n_clusters=2, metric="dtw", max_iter=500, random_state=i*2)
    km.fit(X_train)
    print(silhouette_score(X_train,
                           km.predict(X_train),
                           metric="dtw"))

Results in

0.1379816483983626
0.12501642312525266 # lowest
0.19600175472468492
0.21242253672651584 # highest
0.16945987450892622
0.1696772853353848
0.1924580995443414
0.17290448200720934
0.14888132304396023

Repeat the above procedure with n=3 clusters and print silhouette score:

for i in range(2,11):
    km = TimeSeriesKMeans(n_clusters=3, metric="dtw", max_iter=500, random_state=i*2)
    km.fit(X_train)
    print(silhouette_score(X_train,
                           km.predict(X_train),
                           metric="dtw"))

Results in

0.20512134769791557
0.23869052731302476 # highest
0.1988968698176969
0.210215409064512
0.21938781768333507
0.20512134769791557
0.23869052731302476 # highest
0.17423182636957274
0.10851683338892922 # lowest

So depending on the result, either n=2 or n=3 would be selected, but it is not unambiguous

Expected behavior
I would have expected a a more or less consistent silhouette score (maybe around +- 0.02). At least that the levels stay the same.

Environment (please complete the following information):

  • OS: Windows 10 Enterprices
  • tslearn version 0.4.1

Additional context
I detected the problem in a different data set which I did not share here. There, the differences were even higher (n=2) sometimes returned 0.17 (lowest) and 0.52 (highest)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the provided CBF reproduction with TimeSeriesKMeans and silhouette_score, comparing repeated runs for two and three clusters. Trace the behavior of these two entry points to determine whether the variation is expected or indicates a bug, then make the result consistent or clarify the documented behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.