confluentinc / confluentinc/kafka-tutorials
Clean up old pages from S3
- Dominant language
- Java
- Stars
- 39
- Forks
- 91
- PR merge metrics
- No merged PRs in 30d
Description
https://kafka-tutorials.confluent.io/kafka-connect-datagen-ccloud/kafka.html was removed a few months back by PR https://github.com/confluentinc/kafka-tutorials/pull/856 but it is still a live page.
Why is this happening? I suspect this is because this page persists in S3 and so it keeps getting indexed and discoverable.
This GH issue is two parts:
1. Clean out the S3 bucket and ensure only intended pages are discoverable
2. Handle redirects for bad links, otherwise the current UX is below
Contributor guide
Research direction
Start by reviewing PR #856 and the repository’s page publication process to identify how removed pages remain in S3. Verify which objects are intended to remain, then inspect how bad links are handled. Done means the removed page is no longer discoverable and invalid links have the agreed redirect behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws
- Domain
- cloud, documentation
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100