apache / apache/kyuubi

[Improvement] Using [session user] as [proxy user] to execute statements at the server share level

Open
#5,687 8 comments 0 reactions 0 assignees View on GitHub
Dominant language
Scala
Stars
2.4k
Forks
1k
PR merge metrics
No merged PRs in 30d

Description

### Code of Conduct

- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

### Search before asking

- [X] I have searched in the [issues](https://github.com/apache/kyuubi/issues?q=is%3Aissue) and found no similar issues.

### What would you like to be improved?

**At the server share level, If I don't use the Ranger AuthZ Plugin,** multiple users use the system user's credential, who submits the spark job, to access HDFS, rather than their owner credential.

So on the HDFS side, all actions are executed as the system user, so the HDFS Ranger Plugin loses control.

I want to find a way for all session users can use their own Identity to access HDFS. Thus at the server share level, session users could be under minimum access control.

### How should we improve?

I read Kyuubi's code And found the function setting appropriate configs for each session user. 👇

https://github.com/apache/kyuubi/blob/master/externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/operation/SparkOperation.scala#L130

Maybe we can set [session user] as [proxy user] to execute statements, like this: 👇

![image](https://github.com/apache/kyuubi/assets/6671568/9ab86729-cedb-44b7-bea2-0e709397fc8d)

Are there some lurking issues or concerns at the system level?

Looking forward to your opinions very much!

### Are you willing to submit PR?

- [X] Yes. I would be willing to submit a PR with guidance from the Kyuubi community to improve.
- [ ] No. I cannot submit a PR at this time.

Contributor guide

Open the contributing guide

Research direction

Start with the per-session configuration logic in externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/operation/SparkOperation.scala around line 130. Trace how session and system identities are used for HDFS access, then assess the system-level concerns of using a session user as a proxy user. Done requires a settled approach and explicit implementation scope; the issue names no tests or additional target files.

Written by the indexing model from the issue text.

Assessment

Tech stack
hadoop, scala, spark
Domain
authorization, backend, distributed-systems, security
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.