Nobody Is in Charge and That Is Why It Scales
Last timeSame Model, Different Data
Summing a value across every device looks like it should cost more as devices are added. Arranged as a ring it does not, and the derivation is the most useful thing in this course.
The previous lesson treated the gradient exchange as a cost and measured it.
This lesson opens it up, because how it is arranged changes whether a cluster
can usefully be large.
The obvious way
Pick one machine to be in charge. Every device sends it their gradient, it adds
them up, and it sends the sum back to everyone.
This is correct and it is what anybody draws first. Count the traffic at the
collector. With sixteen devices, it receives fifteen gradients and sends
fifteen back: thirty times the gradient size through one network connection
while the other fifteen machines sit idle. With a hundred and twenty-eight
devices it is two hundred and fifty-four times.
The exchange therefore gets slower as the cluster grows. That is the opposite
of what a cluster is for, and it is why nobody does this.
The ring
Arrange the devices in a circle, each one sending only to its right-hand
neighbour and receiving only from its left. No device talks to any other, and
nobody is in charge.
Split the gradient into as many chunks as there are devices. With four devices,
four chunks, each a quarter of the gradient.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents