microsoft / microsoft/mattergen

Inquiry regarding detailed generation parameters and evaluation metrics for conditional tasks

Open
#244 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
1.8k
Forks
351
Avg merge
1h 3m
Merged PRs (30d)
4

Description

Hello!
I am currently digging into the conditional generation experiments described in your paper, specifically the tasks involving target chemistry, energy_above_hull, and dft_band_gap. While the main text and Appendix D provide a great overview, I am hoping to get a bit more technical detail to help us better understand the pipeline and apply it to our own research in materials informatics.

Could you please provide some guidance or share data on the following?

  1. Detailed Generation Parameters
    Could you share the specific sampling parameters and configurations used for these conditional tasks? For instance, I would love to know the exact configuration format or code snippet used to apply joint conditions, such as {'chemical_system': 'O-Sr-V', 'energy_above_hull': 0.0}.

  2. MatterSim Evaluation and Sampling Details
    Could you clarify the exact sampling sizes used across the different chemical systems (e.g., the specific breakdown for ternary and quaternary systems, alongside the 10,240 mentioned for quinary)? Additionally, are the detailed evaluation results—especially the explicit S.U.N.% metrics for these specific tasks—available, perhaps in a tabular format?

Thank you for your time and for sharing this fantastic work with the community!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with Appendix D and the conditional-generation experiments for target chemistry, energy_above_hull, and dft_band_gap. Determine whether the sampling configurations, ternary and quaternary sample breakdowns, and S.U.N.% evaluation tables exist in the project or paper materials. Done means documenting or linking the requested parameters and evaluation results.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.