dfe-analytical-services / dfe-analytical-services/analysts-guide

Databricks - information on options based on team's current process

Open
#184 2 comments 0 reactions 0 assignees View on GitHub
enhancement good first issue
Dominant language
R
Stars
6
Forks
9
PR merge metrics
No merged PRs in 30d

Description

**Is your feature request related to a problem? Please describe.**
Reading through the Databricks guidance, it seems Teams will be left with the question 'how does this impact my current process?'. I think it would be good for us to have either best practice advice, or what options are available for teams to choose from. For example, for the SQL code, they may choose to adapt for Databricks SQL and run through R, or they may chose to use the Databricks platform.

**Describe the solution you'd like**
I think it would be useful for there to be a table or flow chart which describes what options teams have, based on their current process. For example, what do teams need to do if their code is all R code, what do teams need to do if their code is a mix of SQL/R code, what do teams need to do if they heavily use SQL code. Where the team currently store their code may need to be considered too.

**Describe alternatives you've considered**
Not really any - I think it should be something easy to follow, with links to the relevant Databricks guidance.

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the existing Databricks guidance and the linked guidance it already references. Define a table or flow chart covering R-only, mixed SQL/R, and SQL-heavy processes, including where code is stored; done means each scenario has clear options, actions, and relevant Databricks links.

Written by the indexing model from the issue text.

Assessment

Tech stack
r
Domain
documentation
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.