CsvSource type conversion with custom schema
- Vorherrschende Sprache
- Scala
- Sterne
- 147
- Forks
- 32
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
From the project README - [CSV source part](https://github.com/51zero/eel-sdk#csv-source) I got the idea that type conversion for loaded CSV should be performed according to the specified schema.
But if I define a custom schema for a `CsvSource` which has columns with other types than `String` (`Int` for example), then the values in that column are still returned as `String`.
Is it intended behaviour, bug or it just haven't been implemented?
Runnable example:
```scala
import java.io.ByteArrayInputStream
import java.nio.charset.StandardCharsets
import io.eels.component.csv.CsvSource
import io.eels.schema._
object CsvSourceTypeConversionTest extends App {
val exampleCsvString =
"""A,B,C,D
|1,2.2,3,foo
|4,5.5,6,bar
""".stripMargin
val stream = new ByteArrayInputStream(exampleCsvString.getBytes(StandardCharsets.UTF_8))
val schema = new StructType(Vector(
Field("A", IntType.Signed),
Field("B", DoubleType),
Field("C", IntType.Signed),
Field("D", StringType)
))
val ds = new CsvSource(stream _, Some(schema)).toDataStream()
val firstRow = ds.iterator.toIterable.head
val firstRowA = firstRow.get("A")
println(firstRowA) // prints 1 as expected
println(firstRowA.getClass.getTypeName) // prints java.lang.String
assert(firstRowA == 1) // this assertion will fail because firstRowA is not an Int
}
```
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Rechercherichtung
Look at the CsvSource implementation in the eel-sdk codebase, likely under io.eels.component.csv. Examine how the schema is applied during row parsing. The test case provided shows that values are returned as Strings despite Int/Double schema. Start by running the provided example to confirm the issue, then trace the data flow from CSV parsing to row creation. Check if type conversion logic exists or needs to be added. Verify by writing a test in the project's test suite.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- scala
- Bereich
- data-engineering
- Issue-Typ
- Bug
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Klar beschrieben
- Anfängerfreundlichkeit
- 45/100