export_to_phy vs using kilosort sorter output params.py directly

未关闭
#4,635 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
45/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
python
领域
data

调研方向

Start at the export_to_phy function and compare the files it produces with the params.py and TSV files used by the direct Kilosort4 workflow. Reproduce both workflows on the same sorting, then identify what extra work accounts for the runtime and fewer waveform channels, and document whether the direct workflow is equivalent.

由索引模型根据 Issue 内容生成。

描述

exporters

Hello!
We realised in our lab recently that two of us were using two different ways to manually curate units with Phy after sorting with kilosort4.
The first is to use to use the export_to_phy function and then open the params.py file with Phy as described in the documentation (the exporting is very slow, sometimes slower than the actual sorting):

analyzer = si.create_sorting_analyzer(sorting_KS4, recording_saved, sparse=True)
# compute all the extensions required
sexp.export_to_phy(sorting_analyzer=analyzer, output_folder=Path(data_dir) / 'phy_folder', verbose=True, copy_binary=False)

The second is to just compute the extensions needed, save them as tsv files, copy to the sorter output location where params.py from the sorter output is generated, and then just open that with Phy without the export_to_phy function (much faster, barely any extra compute time).

sorting_analyzer = si.create_sorting_analyzer(sorting=sorting_KS4, recording=rec_corrected, format="binary_folder", folder = KSfolder / 'analyzer_med' )
contamination = sqm.compute_sliding_rp_violations(sorting_analyzer=sorting_analyzer,
                                                  bin_size_ms=0.25)

presence_ratio = sqm.compute_presence_ratios(sorting_analyzer=sorting_analyzer)

def save_dict_to_tsv(data, header_name, file_path, delimiter='\t'):
    """
    Saves a dictionary to a TSV file.

    Args:
        data (dict): The dictionary to save. Keys will be the header row.
        file_path (str): The path to the TSV file.
        delimiter (str, optional): The delimiter. Defaults to tab ('\t').
    """
    #with open(file_path, 'w', newline='', encoding='utf-8') as tsvfile:
    with open(Path(KSfolder) / 'sorter_output' / file_path, 'w', newline='', encoding='utf-8') as tsvfile:

        writer = csv.writer(tsvfile, delimiter=delimiter)

        writer.writerow(['cluster_id', header_name])
        for key, value in data.items():
            writer.writerow([key, value])

save_dict_to_tsv(contamination, 'sliding_rp', 'cluster_sliding_rp.tsv')
save_dict_to_tsv(presence_ratio, 'presence', 'cluster_presence.tsv')
# move these to the same folder that holds your params.py file for phy


We tried both for the same sorting, and the only thing that jumped out to us was that the second method resulted in fewer channels in the waveform view on Phy, but no other noticeable difference.

What does export_to_phy do that takes so much time, and is it necessary to do it, since the second method seems to be working fine? Or are we missing something here?

主要语言
Python
星标
847
派生
280
平均合并
3 天 9 小时
30 天内合并 PR
29

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

SpikeInterface/spikeinterface 的其他 Issue

查看 SpikeInterface/spikeinterface 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。