1
00:00:00,100 --> 00:00:01,540
Welcome back to the Architecture
Corner. 

2
00:00:01,550 --> 00:00:04,110
I'm your host and really glad 
you could join us for Part 2 of 

3
00:00:04,120 --> 00:00:07,070
our five part journey into 
application availability. 

4
00:00:07,140 --> 00:00:10,250
Yeah, it's great to be back. 
So last time we sort of set the 

5
00:00:10,260 --> 00:00:12,630
stage, didn't we? 
Talk about availability itself, 

6
00:00:12,640 --> 00:00:15,650
why segmenting your app is so 
crucial? 

7
00:00:15,660 --> 00:00:17,730
Exactly. 
And today we're getting into the

8
00:00:17,740 --> 00:00:21,000
nuts and bolts. 
Really 2 core ideas you 

9
00:00:21,010 --> 00:00:24,740
absolutely need for keeping 
things running, redundancy and 

10
00:00:24,750 --> 00:00:26,000
persistence. 
Hmm. 

11
00:00:26,050 --> 00:00:28,840
The safety Nets, essentially. 
Right, so let's unpack this. 

12
00:00:28,850 --> 00:00:33,380
We'll look at how having 
backups, backup systems, and 

13
00:00:33,390 --> 00:00:37,540
managing data the right way can 
save you when things inevitably 

14
00:00:37,550 --> 00:00:40,740
go sideways. 
OK, so redundancy first, it's 

15
00:00:40,750 --> 00:00:42,680
hard. 
It's just, well, it's about 

16
00:00:42,690 --> 00:00:44,160
having more than one of 
something. 

17
00:00:44,610 --> 00:00:46,490
Multiple instances of a 
resource. 

18
00:00:46,670 --> 00:00:49,380
Could be compute, storage, 
anything really. 

19
00:00:49,940 --> 00:00:51,190
And breaks. 
Exactly. 

20
00:00:51,200 --> 00:00:54,250
If one resource goes offline, 
traffic just gets rerouted 

21
00:00:54,840 --> 00:00:56,540
pretty quickly to one of its 
siblings. 

22
00:00:56,550 --> 00:00:58,710
Another healthy instance. 
OK, so it's kind of like having 

23
00:00:58,720 --> 00:01:00,990
a backup generator for your 
house, but maybe for your web 

24
00:01:01,000 --> 00:01:03,560
servers instead? 
That's a decent analogy, yeah. 

25
00:01:03,640 --> 00:01:07,570
And a common way to handle this,
I think the article by Mario 

26
00:01:07,580 --> 00:01:10,610
Bittencourt touches on this, is 
using things like reverse 

27
00:01:10,620 --> 00:01:14,880
proxies or load balancers. 
Traffic cops. 

28
00:01:14,930 --> 00:01:18,700
Pretty much the client talks to 
the proxy, and the proxy is 

29
00:01:18,710 --> 00:01:21,340
smart enough to send that 
request only to servers that are

30
00:01:21,350 --> 00:01:24,800
actually up and running. 
In that list of healthy servers 

31
00:01:24,810 --> 00:01:27,260
changes dynamically, right? 
It's not static. 

32
00:01:27,310 --> 00:01:30,280
Oh absolutely, that's key. 
The load balancer is constantly 

33
00:01:30,290 --> 00:01:31,760
checking. 
If a server fails a health 

34
00:01:31,770 --> 00:01:33,940
check, boom, stops getting 
traffic. 

35
00:01:34,010 --> 00:01:37,400
This happens automatically even 
in say, Kubernetes environments.

36
00:01:37,650 --> 00:01:41,100
Requests just flow between 
healthy pods as they spin up or 

37
00:01:41,110 --> 00:01:42,610
down. 
OK, but here's where it gets a 

38
00:01:42,620 --> 00:01:45,690
bit tricky. 
I think redundancy isn't always 

39
00:01:45,700 --> 00:01:48,760
straightforward, especially if 
your application isn't 

40
00:01:48,770 --> 00:01:51,000
stateless. 
Right, that's a huge factor. 

41
00:01:51,050 --> 00:01:53,960
If your app needs to remember 
things between requests, like 

42
00:01:53,970 --> 00:01:57,880
user sessions or stores 
temporary data locally, well, 

43
00:01:57,950 --> 00:02:00,900
switching to a totally new 
instance can cause problems. 

44
00:02:00,910 --> 00:02:03,290
Because the new instance doesn't
have that context. 

45
00:02:03,390 --> 00:02:07,160
Exactly that session info, that 
local data, it isn't shared 

46
00:02:07,170 --> 00:02:09,780
automatically. 
Stateless services much easier. 

47
00:02:09,850 --> 00:02:12,760
They don't store that local 
context, so any instance can 

48
00:02:12,770 --> 00:02:15,970
handle any request. 
It makes redundancy simpler to 

49
00:02:15,980 --> 00:02:17,150
implement. 
Got it. 

50
00:02:17,220 --> 00:02:20,710
So that's compute redundancy, 
but what about the data itself? 

51
00:02:20,720 --> 00:02:22,650
Well, that brings us nicely to 
persistence. 

52
00:02:22,950 --> 00:02:27,100
See, most applications have the 
code, the logic, and then the 

53
00:02:27,110 --> 00:02:30,050
data it works with. 
The code might not change often,

54
00:02:30,220 --> 00:02:34,420
but the data does constantly OK,
so making sure your data is 

55
00:02:34,430 --> 00:02:37,040
redundant. 
Achieving persistence has its 

56
00:02:37,050 --> 00:02:40,760
own set of challenges. 
A single change to the data can 

57
00:02:40,770 --> 00:02:42,740
make different copies instantly 
out of sync. 

58
00:02:43,430 --> 00:02:45,990
Right, so how do you keep all 
those copies consistent if you 

59
00:02:46,000 --> 00:02:48,200
have multiple databases, say? 
Good question. 

60
00:02:48,470 --> 00:02:51,340
The main approach is boil down 
to two types of replication, 

61
00:02:51,770 --> 00:02:54,490
synchronous and asynchronous. 
OK, sync and async. 

62
00:02:54,500 --> 00:02:55,320
What's the? 
Difference. 

63
00:02:55,330 --> 00:02:57,680
Well, with synchronous 
replication, a piece of data 

64
00:02:57,690 --> 00:03:01,160
isn't considered saved until 
multiple servers. 

65
00:03:01,170 --> 00:03:03,850
The primary end replicas confirm
they've. 

66
00:03:03,860 --> 00:03:05,460
Got it. 
So it waits for everyone. 

67
00:03:05,530 --> 00:03:08,090
Kind of, yeah. 
It waits for confirmation from 

68
00:03:08,100 --> 00:03:10,380
multiple places. 
The big advantage you get 

69
00:03:10,390 --> 00:03:14,420
guarantees no data loss and 
strong consistency. 

70
00:03:14,430 --> 00:03:17,020
Everyone sees the same data. 
But the downside is. 

71
00:03:17,170 --> 00:03:19,850
Latency. 
Because you're waiting for those

72
00:03:19,860 --> 00:03:24,080
confirmations. 
Every right takes longer and 

73
00:03:24,440 --> 00:03:26,990
handling failures gets a bit 
more complex too. 

74
00:03:27,140 --> 00:03:30,610
OK, So what about asynchronous? 
Asynchronous is different. 

75
00:03:30,760 --> 00:03:34,950
Data is marked saved as soon as 
the primary server gets it, then

76
00:03:35,000 --> 00:03:37,950
replication of the other servers
happens while in the background.

77
00:03:38,260 --> 00:03:40,470
So it feels faster to the user 
initially a. 

78
00:03:40,480 --> 00:03:43,030
Definitely lower latency and 
it's generally simpler to 

79
00:03:43,040 --> 00:03:45,730
manage. 
But, and this is important for 

80
00:03:45,740 --> 00:03:49,410
you, comes with tradeoffs. 
Big ones like what potential 

81
00:03:49,420 --> 00:03:52,590
data loss if the primary server 
fails before it manages to 

82
00:03:52,600 --> 00:03:55,760
replicate the data will that 
data might be gone out. 

83
00:03:55,880 --> 00:03:58,730
And the other issue is you might
serve stale data. 

84
00:03:59,040 --> 00:04:01,450
If you read from a replica that 
hasn't received the latest 

85
00:04:01,460 --> 00:04:03,860
update yet, you could see old 
information. 

86
00:04:03,870 --> 00:04:07,070
That's the classic read after 
write consistency issue. 

87
00:04:07,080 --> 00:04:09,630
Right, so you update something 
immediately, read it back and 

88
00:04:09,640 --> 00:04:11,270
you might get the old version. 
Exactly. 

89
00:04:11,280 --> 00:04:14,760
So how does this play out in the
real world like of cloud 

90
00:04:14,770 --> 00:04:18,399
providers, AWS for example? 
Yeah, AWS provides good 

91
00:04:18,410 --> 00:04:20,500
examples. 
They heavily use their 

92
00:04:20,510 --> 00:04:26,820
availability zone AZZ for this. 
Think of Aziz as separate data 

93
00:04:26,830 --> 00:04:29,650
centres, physically isolated but
connected with really fast 

94
00:04:29,660 --> 00:04:31,670
networks. 
They're the building blocks for 

95
00:04:31,680 --> 00:04:33,830
redundancy. 
And different services use these

96
00:04:33,840 --> 00:04:37,400
AZ differently for persistence. 
They do take RDS, their 

97
00:04:37,410 --> 00:04:40,160
relational database service. 
You can set up synchronous 

98
00:04:40,170 --> 00:04:42,720
replication to a replica and a 
different azz. 

99
00:04:43,210 --> 00:04:47,970
That gives you high up time, 
maybe 99.95%, but failover if 

100
00:04:47,980 --> 00:04:51,410
the primary dies might take say 
30 to 60 seconds. 

101
00:04:51,600 --> 00:04:53,130
OK. 
Still some downtime potential 

102
00:04:53,140 --> 00:04:54,010
there the. 
Little yeah. 

103
00:04:54,020 --> 00:04:56,170
Then there's Aurora. 
It's designed differently. 

104
00:04:56,280 --> 00:04:58,990
Data is automatically copied 
across 6 storage nodes and three

105
00:04:59,000 --> 00:05:00,740
AZ. 
Even if you don't set up read 

106
00:05:00,750 --> 00:05:04,290
replicas yourself, it aims for 
super low replication lag, often

107
00:05:04,300 --> 00:05:06,810
under 100 milliseconds. 
Wow, that's fast, but still 

108
00:05:06,820 --> 00:05:08,410
asynchronous. 
Fundamentally, yes. 

109
00:05:08,460 --> 00:05:11,250
So there's still that tiny, tiny
chance of reading slightly stale

110
00:05:11,260 --> 00:05:14,170
data from a replica, though they
work hard to minimize it. 

111
00:05:14,240 --> 00:05:16,370
Interesting. 
And Dynamo DB that's different 

112
00:05:16,380 --> 00:05:18,070
again isn't. 
It very different. 

113
00:05:18,380 --> 00:05:21,230
Dynamodb uses a really 
interesting hybrid approach. 

114
00:05:21,620 --> 00:05:24,930
When you write data, it's 
synchronously copied to two 

115
00:05:24,940 --> 00:05:27,810
nodes in different azz, ensuring
durability. 

116
00:05:28,160 --> 00:05:31,870
Then it's asynchronously copied 
to 1/3 node in another azz. 

117
00:05:32,720 --> 00:05:36,090
So you get the safety of three 
copies across AZ, but you only 

118
00:05:36,100 --> 00:05:38,650
wait for two confirmations. 
Exactly. 

119
00:05:38,660 --> 00:05:41,430
It's a balance between 
durability, availability and 

120
00:05:41,440 --> 00:05:43,910
latency. 
And then when you read data from

121
00:05:43,920 --> 00:05:46,930
Dynamodb, you the user get a 
choice. 

122
00:05:46,940 --> 00:05:48,800
A choice. 
Yep, you can ask for an 

123
00:05:48,810 --> 00:05:51,970
eventually consistent read, 
which is faster but might return

124
00:05:51,980 --> 00:05:54,730
stale data. 
Or you can request a strongly 

125
00:05:54,740 --> 00:05:58,110
consistent read which guarantees
the absolute latest data, but 

126
00:05:58,120 --> 00:06:00,530
might be slightly slower because
it always goes to the primary 

127
00:06:00,540 --> 00:06:03,170
node cluster. 
So you tailor the consistency to

128
00:06:03,180 --> 00:06:05,940
your application specific need 
for that particular read. 

129
00:06:05,950 --> 00:06:08,150
Precisely, it puts the trade off
decision in your hands. 

130
00:06:08,160 --> 00:06:11,770
OK, so wrapping this up, it 
seems like redundancy, whether 

131
00:06:11,780 --> 00:06:15,370
for compute or data persistence 
is fundamental for availability,

132
00:06:15,420 --> 00:06:17,280
but it's never free. 
Never. 

133
00:06:17,550 --> 00:06:20,580
You're always playing with 
trade-offs, the safety of having

134
00:06:20,590 --> 00:06:24,010
multiple copies versus 
potentially higher latency, or 

135
00:06:24,020 --> 00:06:26,340
the complexities of eventual 
consistency. 

136
00:06:26,430 --> 00:06:29,840
Which leads to a really critical
question for you listening How 

137
00:06:29,850 --> 00:06:32,980
do you actually decide what 
level of consistency or what 

138
00:06:32,990 --> 00:06:36,360
latency is genuinely good enough
for your application? 

139
00:06:36,430 --> 00:06:39,540
Yeah, that's not always obvious.
It depends entirely on the 

140
00:06:39,550 --> 00:06:42,460
specific requirements and 
failure modes you're willing to 

141
00:06:42,470 --> 00:06:45,450
tolerate. 
It's definitely not A1 size fits

142
00:06:45,460 --> 00:06:47,260
all situation. 
Absolutely. 

143
00:06:47,420 --> 00:06:50,000
And there's still more to cover.
We've really just scratched the 

144
00:06:50,010 --> 00:06:52,720
surface here. 
Join us next time as we continue

145
00:06:52,730 --> 00:06:56,270
exploring availability, looking 
at concepts like graceful 

146
00:06:56,280 --> 00:06:59,200
degradation and asynchronous 
processing. 

147
00:06:59,250 --> 00:07:01,620
More tools for your resilience 
toolkit. 

148
00:07:01,670 --> 00:07:04,800
Definitely, And for more details
on today's topics and some 

149
00:07:04,810 --> 00:07:07,800
deeper dives, do check out the 
description for this discussion.

150
00:07:07,910 --> 00:07:10,700
And please don't forget to 
subscribe for free to the 

151
00:07:10,710 --> 00:07:12,200
Architecture Corner newsletter 
over at 

152
00:07:12,210 --> 00:07:15,460
architecturecorner.substack.com.
Thanks so much for joining us in

153
00:07:15,470 --> 00:07:17,360
the Architecture Corner. 
We'll see you next time.

