Allow sources to be configured as 'always have data'
- Ngôn ngữ chính
- Scala
- Star
- 31
- Fork
- 4
- Merge trung bình
- 1 ngày 10 phút
- Pull request đã merge (30 ngày)
- 4
Mô tả
## Background
Currently, ingestion jobs always checks the record count first, and only if the record count is non-zero, proceeds with the ingestion. This works well for event-based JDBC sources. But for file based sources it sometimes requires to parse data twice: the first time to get the record count, and the second time to load the data.
## Feature
Extend the interface of `Source` to allow instances where data is always assumed to be available, like snapshot sources.
When a source has 'always data available' the record count should be derived on write.
## Things to consider
- If the data does not get through the metastore, but passed directly from a source to a sink, possibly requires caching data, or saving in a temporary directory.
## Proposed Solution
```scala
trait Source extends ExternalChannel {
/**
* If true, getRecordCount() won't be used to determine if the data is available.
* This saves performance ono double data read when the source is file based and always points to a particular file,
* or it is a snapshot-based source
*/
def isDataAlwaysAvailable: Boolean = false
}
```
The default implementation is `false` so that the interface is source code compatible with existing sources.
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Đánh giá
Issue này chưa được đánh giá.