deepseek-ai / deepseek-ai/DeepSeek-Math

how to sample 64 output from old policy model?

Open
#14 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.4k
Forks
592
PR merge metrics
No merged PRs in 30d

Description

Is it just adjusting the decoding parameters?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no file, test, or entry point and does not identify the old policy model or sampling command. Start by clarifying which model and inference procedure are intended, then verify whether producing 64 outputs depends on decoding parameters and document the confirmed usage.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.