Handling failure in a distributed stream processing system is challenging, but Flink removes a lot of this complexity from the job of the application developer.
A high-availability Flink cluster will often use an Apache Zookeeper cluster to elect a leader Job Manager (coordinator) instance. One or more Task Managers (the systems that actually execute the pieces of a Flink application) discover the current Job Manager leader by querying Zookeeper.
Zookeeper tracks the current leader and running jobs. State for those jobs is written periodically to external storage, often an HDFS cluster. Flink automatic-checkpointing will write application state at fixed intervals to storage and, in the event of a failure, automatically resume from the most recent checkpoint. There's also support for manual savepoints which can be used to restore state when submitting a new job or resuming from catastrophic failure.
Flink provides exactly-once guarantees within the context of the Flink application; any side-effects of your application, such as calls to external services or records written to a database, can happen multiple times if you're recovering from a failure.