h4cc / h4cc/slugger

The “unwanted characters regex” matches wanted characters

Open
#35 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Elixir
Stars
160
Forks
26
PR merge metrics
No merged PRs in 30d

Description

Hi!

I found a potential bug in the following line:

```elixir
|> remove_unwanted_chars(separator, ~r/([^A-Za-z0-9가-힣])+/)
```

The line replaces all characters that does not match `A-Za-z0-9가-힣` with the separator character.

However, we found an unwanted characters that fall under this expression, namely ` ` ([U+2009](https://www.fileformat.info/info/unicode/char/2009/index.htm)).

```elixir
source = "foo bar" # This is "foobar"

String.replace(source, ~r/([^a-z0-9가-힣])+/, "-")
# => "foo bar"

String.replace(source, ~r/([^a-z0-9])+/, "-")
# => "foo-bar"
```

The behavior is the same for the U+2010, U+2011, U+2012, etc. characters.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.