haskell / haskell/text

Add efficient findCharIndex

Open
#445 2 comments 0 reactions 0 assignees View on GitHub
feature request
Dominant language
Haskell
Stars
421
Forks
163
PR merge metrics
No merged PRs in 30d

Description

```haskell
findCharIndex :: Char -> Text -> Maybe Int
```
When having look at #369, it looks like having access to this primitive, with a fast C implementation that uses the known size of each char in utf-8 might be able to make `isSubsequenceOf` quite fast, particularly when the haystack is quite large and the needle chars are spread far apart.

```C
int findCharIndex(const char * haystack, int length, int target) {
char * offset;
if (target < 128) {
offset = memchr(haystack, length target);
} else {
// Implementation left as an exercise to the reader, but it would look for
// utf-8 encoded bytes for the Char.
offset = efficientlyFindUTF8Codepoint(haystack, length, target);
}
if (offset == NULL) return -1;
return (int)(offset - haystack);
}
```

Having this also means we could add the rule

```haskell
{-# RULES "findIndex Char" findIndex (c ==) = findCharIndex c #-}
{-# RULES "findIndex Char2" findIndex (== c) = findCharIndex c #-}
```

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.