benoitc / benoitc/hackney

Connections can get stuck when upgrading to SSL

Open
#916 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Erlang
Stars
1.4k
Forks
455
Avg merge
2d 18h
Merged PRs (30d)
1

Description

### Environment
- hackney: 4.7.2
- Erlang/OTP: 26

### Description
I observed a behavior in an integration environment where the default connection pool becomes fully occupied by connections which were stuck attempting to upgrade to SSL. Within [hackney_con:connected](https://github.com/benoitc/hackney/blob/9d684fb768837a74c4c7beff0b0e024eb9a9998a/src/hackney_conn.erl#L929) there is a call to `ssl:connect/2` which results in utilizing the default timeout of infinity. If an issue occurs during the TLS handshake, the connection process remains in a waiting state and attempts to utilize the connection time out repeatedly as the call to `hackney:maybe_register_h2/6` must pull the protocol from the state of the pooled connection which times out.

Is it possible to implement a concept of tls upgrade timeout or tls handshake timeout which can be optionally passed to an `ssl:connect/3` call in the upgrade to TLS path?

### More Details
Hackney Connection Supervisor has 4 linked processes in the pool and appears in a healthy state
```
[{registered_name,hackney_conn_sup},
{current_function,{gen_server,loop,7}},
{initial_call,{proc_lib,init_p,5}},
{status,waiting},
{message_queue_len,0},
{links,[<0.663122.0>,<0.663147.0>,<0.663279.0>,<0.663134.0>,
<0.1568.0>]},
{dictionary,[{'$ancestors',[hackney_sup,<0.1567.0>]},
{'$initial_call',{supervisor,hackney_conn_sup,1}}]},
{trap_exit,true},
{error_handler,error_handler},
{priority,normal},
{group_leader,<0.1566.0>},
{total_heap_size,2208},
{heap_size,1598},
{stack_size,11},
{reductions,865131},
{garbage_collection,[{max_heap_size,#{error_logger => true,include_shared_binaries => false,
kill => true,size => 0}},
{min_bin_vheap_size,46422},
{min_heap_size,233},
{fullsweep_after,65535},
{minor_gcs,35}]},
{suspending,[]}]
```

One of the connections in the pool is inspected and is found in the waiting state with two messages in the queue. Connection is waiting on ssl handshake to complete
```
23> process_info(<0.663122.0>).
[{current_function,{gen,do_call,4}},
{initial_call,{proc_lib,init_p,5}},
{status,waiting},
{message_queue_len,2},
{links,[<0.1570.0>]},
{dictionary,[{'$ancestors',[hackney_conn_sup,hackney_sup,
<0.1567.0>]},
{'$initial_call',{hackney_conn,init,1}}]},
{trap_exit,true},
{error_handler,error_handler},
{priority,normal},
{group_leader,<0.1566.0>},
{total_heap_size,3203},
{heap_size,1598},
{stack_size,47},
{reductions,14839},
{garbage_collection,[{max_heap_size,#{error_logger => true,include_shared_binaries => false,
kill => true,size => 0}},
{min_bin_vheap_size,46422},
{min_heap_size,233},
{fullsweep_after,65535},
{minor_gcs,9}]},
{suspending,[]}]

26> process_info(<0.663122.0>, [current_stacktrace]).
[{current_stacktrace,[{gen,do_call,4,
[{file,"gen.erl"},{line,240}]},
{gen_statem,call,3,[{file,"gen_statem.erl"},{line,923}]},
{ssl_gen_statem,call,2,
[{file,"ssl_gen_statem.erl"},{line,1372}]},
{ssl_gen_statem,handshake,2,
[{file,"ssl_gen_statem.erl"},{line,253}]},
{ssl_gen_statem,connect,8,
[{file,"ssl_gen_statem.erl"},{line,222}]},
{ssl,connect,3,[{file,"ssl.erl"},{line,637}]},
{hackney_conn,connected,3,
[{file,"/codebuild/output/src3735438855/src/mirror/_build/default/lib/hackney/src/hackney_conn.erl"},
{line,909}]},
{gen_statem,loop_state_callback,11,
[{file,"gen_statem.erl"},{line,1395}]}]}]
```

At this point in time I am not exactly sure what failure occurred to cause the handshake attempt to fail in this way so I can not reliably reproduce this issue. Regardless of the failure scenario I believe a timeout in this location would allow configuration of some safety rails against whatever kind of failures could happen.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.