[FEA]: Generalize autotuning to dynamically generated kernels
Nessuno ha ancora preso questa issue.
- Lingua principale
- Python
- Stelle
- 2.2k
- Fork
- 155
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Is this a new feature, an improvement, or a change to existing functionality?
Improvement
How would you describe the priority of this feature request?
Low (would be nice)
Please provide a clear description of problem this feature solves
Currently, exhaustive_search takes a single fixed kernel and construct arguments from an arbitrary sequence of configurations. That it takes a fixed kernel makes it unusable in its current form if the config objects themselves generates the kernel.
Feature Description
I'm currently using cutile and metaprogramming to generate kernels, with much better success and less pain than other frameworks. However, I can't use exhaustive_search in its current form, but need to modify it to take kernel generating functions.
Describe your ideal solution
The proposal is essentially to change exhaustive_search, or add a separate case, where we generate the kernel candidate from a config, i.e:
...
for i, cfg in enumerate(search_space):
if not quiet and isatty:
progress(0, i, total, len(errors))
grid = grid_fn(cfg)
kernel = kernel_fn(cfg)
hints = hints_fn(cfg) if hints_fn is not None else {}
updated_kernel = kernel.replace_hints(**hints)
candidate = _TimingCandidate(
config=cfg,
grid=grid,
kernel=updated_kernel,
get_args=lambda _cfg=cfg: args_fn(_cfg),
)
...
Describe any alternatives you have considered
No response
Additional context
No response
Contributing Guidelines
- I agree to follow cuTile Python's contributing guidelines
- I have searched the open feature requests and have found no duplicates for this feature request
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia leggendo l’entry point exhaustive_search e l’utilizzo di _TimingCandidate descritti nell’issue, quindi segui il modo in cui le configurazioni producono attualmente kernel, grid, hint e argomenti. Il lavoro è completo quando exhaustive_search può supportare kernel e argomenti generati da ogni configurazione, preservando il comportamento esistente dei kernel fissi e i candidati di autotuning.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- backend, performance
- Tipo di issue
- Funzionalità
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Tranquilla
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 45/100