althonos / althonos/pyhmmer

Question about extract columns from hmmsearch domain output

未关闭
#82 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
question
主要语言
Cython
星标
170
派生
20
平均合并
58 分钟
30 天内合并 PR
1

描述

Hi,

Thank you for designing this program! I'm currently using it to replace the HMMER in my python script, but I meet some issues currently:

available_memory = psutil.virtual_memory().available
target_size = os.stat(self.input_faa).st_size
hmm_files=HMMFiles(self.hmm_file)
with open("test_hmmer_result.txt", 'wb') as f:
with pyhmmer.easel.SequenceFile(self.input_faa, digital=True) as seqs:
if target_size < available_memory * 0.1: #Pre-fetching targets into memory
targets = seqs.read_block()
else:
targets = seqs
for i, hits in enumerate(pyhmmer.hmmsearch(hmm_files, targets, cpus=os.cpu_count(), domE=1e-15)):
hits.write(f, format="domains", header=False)

I'm using hits.write(f, format="domains", header=False) to get the domain output, but I want to extract columns with [query name qlen target name tlen i-Evalue(this domain) hmm_from ali_from] to form a new output result with tab as a separator. I read the doc file but I'm still confused about how to extract those information.

Could you please tell me how Tophit was selected? I noticed that I may have two Tophits with one protein accession from the hits.write.

Really appreciate your help!

Best Regards,
XInpeng

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。