dfe-analytical-services / dfe-analytical-services/analysts-guide
Databricks - information on options based on team's current process
- Dominant language
- R
- Stars
- 6
- Forks
- 9
- PR merge metrics
- No merged PRs in 30d
Description
**Is your feature request related to a problem? Please describe.**
Reading through the Databricks guidance, it seems Teams will be left with the question 'how does this impact my current process?'. I think it would be good for us to have either best practice advice, or what options are available for teams to choose from. For example, for the SQL code, they may choose to adapt for Databricks SQL and run through R, or they may chose to use the Databricks platform.
**Describe the solution you'd like**
I think it would be useful for there to be a table or flow chart which describes what options teams have, based on their current process. For example, what do teams need to do if their code is all R code, what do teams need to do if their code is a mix of SQL/R code, what do teams need to do if they heavily use SQL code. Where the team currently store their code may need to be considered too.
**Describe alternatives you've considered**
Not really any - I think it should be something easy to follow, with links to the relevant Databricks guidance.
Contributor guide
Research direction
Start by reviewing the existing Databricks guidance and the linked guidance it already references. Define a table or flow chart covering R-only, mixed SQL/R, and SQL-heavy processes, including where code is stored; done means each scenario has clear options, actions, and relevant Databricks links.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- r
- Domain
- documentation
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100