algolia / algolia/algoliasearch-client-java
Adding a configuration for max retries to handle system wide outages
- Lingua principale
- Java
- Stelle
- 52
- Fork
- 33
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
- Algolia Client Version: 3.16.6
- Language Version: Java
### Description
Correct me if I am mistaken but [HttpTransport.executeWithRetry](https://github.com/algolia/algoliasearch-client-java/blob/master/algoliasearch-core/src/main/java/com/algolia/search/HttpTransport.java#L150-L192) will before sending a request, increase the timeout for the request based on the number of retries. After the response is received, it uses *RetryStrategy* to determine what to do with the results.
The function [RetryStrategy.decide](https://github.com/algolia/algoliasearch-client-java/blob/master/algoliasearch-core/src/main/java/com/algolia/search/RetryStrategy.java#L52-L72C4) will in the case of a request time out, will *always* retry.
During an Indexer service outage in July 13th, it appears that was happening for the entire length of the outage causing the Promises to hang and never fail. This becomes a pretty big issue when using `BatchIndexingResponse.waitTask()` like in the API examples.
### Proposed change
1. Adding a new configurable property to ConfigBase called something along the lines of `maxRetriesPerHost` and update `RetryStrategy` to save that value in its constructor.
2. Update `RetryStrategy.decide` to something along the lines of
```
} else if (response.isTimedOut()) {
boolean shouldKeepUp = maxRetriesPerHost == null || tryableHost.getRetryCount() <= maxRetriesPerHost;
tryableHost.setUp(shouldKeepUp);
tryableHost.setLastUse(AlgoliaUtils.nowUTC());
tryableHost.incrementRetryCount();
return RetryOutcome.RETRY;
}
```
### Why make the change
During the time of the indexing outage, it was not clear immediately why our requests were stuck. Looking at the code it seems like if the servers actually responded with server errors, the `StatefulHost`s would have been turned off one by one which would be closer to what I would expect from an outage. The current logic has no way of determining this type of issue, causing ever expanding timeouts without any way of intervening.
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Direzione di ricerca
Inizia con algoliasearch-core/src/main/java/com/algolia/search/HttpTransport.java, in particolare con executeWithRetry, poi leggi RetryStrategy.java e ConfigBase per tracciare la gestione del timeout e del numero di tentativi. Conferma il comportamento durante le richieste con timeout scaduto e considera il lavoro completato quando esiste un limite configurabile di tentativi per host che impedisca ai tentativi di aumentare indefinitamente, preservando il comportamento attuale dei tentativi quando non è impostato.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- java
- Ambito
- api
- Tipo di issue
- Funzionalità
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Ferma
- Chiarezza
- Specificata chiaramente
- Idoneità per principianti
- 45/100