memristors · kt-ram · ai-hardware · supervised-learning · classifier · thermodynamic-computing · generative · sampling · emulator · open-source

Chapter 6: Classification and Thermal Sampling on kT-RAM Neural Lanes

We train neural lanes to classify labelled data. They learn one example at a time and reach the same accuracy as logistic regression. Then we train fresh lanes on the opposite mapping and read them at temperature: clamp a label, draw a sample.

By Alex Nugent ·

Contents
  1. One lane per label
  2. Encoding data as AATs
  3. Fixed bins
  4. Adaptive bins
  5. Adding a bias
  6. Adapt, then freeze
  7. Supervised Learning
  8. The winner is the prediction
  9. Technical abstraction layers
  10. The rank-cut in hardware
  11. Does it work?
  12. Where the misses land
  13. Reading at temperature
  14. Soft and hard feedback
  15. Label in, pattern out
  16. Where we are

At the end of Chapter 5b I said the next move was to hand the lane an answer key, so we show it labelled examples now and have it learn to name them — the simplest supervised task there is.

The task is Iris, the “hello, world” of classifier tasks: a small table of measurements the statistician Ronald Fisher published in 1936, tested by classifiers ever since. It holds a hundred and fifty flowers, fifty from each of three iris species — setosa, versicolor, and virginica — every flower carrying four numbers in centimeters: the length and the width of a petal, and the same two for a sepal, the green leaf-like part beneath the petals. The species is the label, and the job is to read the four numbers and name the flower, small enough to plot on one page and read with your own eyes.

We turn the raw measurements into AATs, feed them into a stack of neural lanes, and train those lanes with a kT-RAM instruction-set routine that gets us logistic regression.

Terminal window
pip install "git+https://github.com/knowm/ktram-neural-core.git#subdirectory=python"

The code example is in examples/iris-classifier in the repo. The whole chapter runs as a workbook you can open in Colab and poke at, covering the encoders, the three-case rule at L0, RankCut, the benchmark, the hot reads, and the chained-against-synchronous sampler:

Open the Iris classifier workbook in Colab

One lane per label#

A single neural lane draws one linear cut through its inputs and answers positive (+) or negative (−). One cut sorts the world into two piles, which is enough for is this a setosa or not, but the Iris dataset has three species. So we use three lanes, one per class, and ask each the yes-or-no question it can actually answer:

  • lane 0 — is this a setosa?
  • lane 1 — is this a versicolor?
  • lane 2 — is this a virginica?

Each lane is its own linear neuron with its own weights, reading the same input, and to classify a flower you read all three and take whichever returns the highest voltage.

Encoding data as AATs#

A lane does not read numbers — it reads AATs, so every input has to become one or more address tuples before a lane sees it. We use “AAT Encoders”, or just “Encoders”, to do this.

The encoder interface is small on purpose: an encoder needs two things. encode(value) returns the AAT, one channel index per space, and space_sizes says how many channels each space holds, which is what the hardware or emulator needs to provision itself. An adaptive encoder adds a third, encode_adapt(value), which tunes its internal state to the data before encoding; one that does not adapt only ever runs encode.

How data gets turned into AATs is a design choice with no single right answer: it depends on the native data type, the resources you have, and what you are trying to do. We will have a lot more to say about it in future chapters — the encoder I’m showing here is a ‘toy’ example, and far more powerful methods exist.

Some data is already most of the way there: an AAT is just integers, and plenty of data is already integers — a category id, a pixel value, or a count. You could take such an integer and feed it in as a channel index unchanged, but usually you should not, because a bare index throws away what the number meant: send 123,456 in as channel 123,456 and the lane learns nothing about its size or its nearness to 123,457, and the two land on unrelated synapses that happen to sit side by side. You could bin the number instead, rounding to the nearest hundred, say, or give each decimal digit its own ten-channel space, so 123,457 encodes to the six-entry AAT (1, 2, 3, 4, 5, 7) — there’s more than one way to skin an AAT, and we don’t need to settle on one here.

A floating point number is not an AAT, so we must encode it: no synapse carries the number 4.2. To get an AAT from it you can bin it such that the bin index is the channel, and since every one of the four Iris measurements is a float, we bin the raw data to create AATs before handing it off to the lanes.

One encoder for this job is the A2DEncoder. The name is analog-to-digital, the same conversion a sensor does when it turns a voltage into a number: it takes a continuous value and reports which bin it landed in, and that bin index is the channel. The name is only an analogy, though: the measurements are already digital numbers, and an AAT is just a different digital encoding of the same information — the encoder is ordinary arithmetic deciding which bin a float belongs to. Each feature gets its own space of bins, sized by a bits knob — bits=3 cuts the range into eight bins.

Fixed bins#

Never let the encoder adapt and the bins stay where they start, an even slicing of [init_min, init_max]:

from ktram_neural_core.encode import A2DEncoder
fixed = A2DEncoder(dims=4, bits=3, init_min=X_tr.min(0), init_max=X_tr.max(0))
# bits=3 -> 8 channels per feature; encode([5.1, 3.5, 1.4, 0.2]) -> (1, 5, 0, 0)

Eight fixed bins per feature already classify Iris about as well as anything does, because Iris is small and nearly separable . But even slices spend bins where no data falls: a bin in an empty stretch never fires, and the flowers all pile into the few bins that cover the crowded stretch. On easy data that waste costs nothing; on skewed or clumpy data it costs a lot.

Adaptive bins#

The A2D encoder can also adapt, moving its grid to match the data distribution: inside it is a binary tree of split points that start at the even slicing and then migrate, every example tugging the nearest edges a little in its direction . Where the data crowds, the bins bunch up and get fine; where it is empty, they stretch and go coarse, until every bin holds about the same number of points — equal-occupancy, with the finest resolution always landing wherever the data is densest.

It is easiest to see on a made-up two-dimensional spread with a few clumps in it, run at a finer grid than Iris needs. Watch the lines crawl off the empty ground and pack into the clusters:

An animated scatter of synthetic two-dimensional data with five blobs of points, overlaid with a grey grid of thirty-two vertical and thirty-two horizontal bin lines. The grid starts evenly spaced, then the lines migrate so they crowd densely through each blob and thin out to wide gaps across the empty space between blobs, until every strip holds roughly the same number of points.
The A2D encoder adapting on a synthetic clumpy spread, thirty-two bins per axis, well past what Iris asks for. The grid starts uniform and migrates toward equal-occupancy: the lines pile up inside each clump, where a small move in the value should change the bin, and stretch across the empty gaps, where it should not. Each axis bins independently, so this is the same one-dimensional adaptation running on x and y at once.

Run the same thing on the actual Iris measurements:

An animated scatter of Iris flowers in petal-length / petal-width space, the three species in blue, green, and red. A grey grid of vertical and horizontal bin lines starts evenly spaced and then migrates, the lines crowding together inside the dense clumps of points and spreading apart across the empty regions, until each strip holds about the same number of flowers.
The A2D encoder adapting its bins to the Iris petal measurements. The grid starts at an even slicing of each axis and then migrates as it walks the data, the lines bunching up where the flowers crowd and stretching across the gaps. By the end every strip holds about the same count, so the encoding spends its resolution where the classes actually sit instead of on empty space.

Adaptive or fixed, each feature still encodes to one active bin, one synapse switched on in its space, and an A2D over the four features produces a four-entry AAT either way. Adaptation moves only the bin edges, not which spaces exist or how many channels they hold, so if you know the distribution, you can compute the bin edges directly and skip the adaptation altogether.

Adding a bias#

Whether the lanes need a bias at all depends on the encoding. Back in Chapter 5 a bias-free node could only draw boundaries through the origin, which is why the overlapping two-synapse encoding there could not reach every state. The same chapter showed the way out: spread the inputs into a balanced encoding, every pattern lighting the same number of channels, and the A2D gives exactly that, one active bin per space, so the encoding itself supplies the freedom a bias would. A separate bias is often unnecessary here; when you do want one, always-on inputs do the job — the ConstantEncoder ignores the value, always lights the same channel, and does not adapt. Stack it onto the A2D with compose, which lays their AATs end to end:

from ktram_neural_core.encode import A2DEncoder, ConstantEncoder, compose
encoder = compose(
A2DEncoder(dims=4, bits=3, init_min=X_tr.min(0), init_max=X_tr.max(0), l=0.01),
ConstantEncoder(), # one always-on synapse — the bias
)
# encode([5.1, 3.5, 1.4, 0.2]) -> (b0, b1, b2, b3, 0); space_sizes [8, 8, 8, 8, 1]

That gives five spaces: four adaptive bins, then the always-on bias channel. The Iris example keeps the bias even though the balanced encoding does not need it, since one extra always-on synapse costs little and shows the mechanism. init_min/init_max seed each feature’s starting grid from the training range, so features on different scales, sepal length against petal width, each get their own bins with no separate normalization step. The composed AAT is the only thing the lanes ever see, and we hand it to the classifier for the rest of the chapter.

Adapt, then freeze#

Adaptive bins move while they settle, so a moving encoding keeps changing the input distribution the lane is trying to fit on every example. It’s usually best to lock the encoding to keep the representation stable , so the training runs in two phases, in order.

First, with the classifier switched off, walk the training data through the encoder and let the bins migrate until they settle; then freeze the encoder and train the lanes against that fixed encoding. The gif above is the first phase — once it stops, the grid holds still, and every flower lands in the same bins every time.

clf.fit(X_tr, y_tr, epochs=5, encoder_epochs=5)

encoder_epochs=5 is phase one — five passes adapting the bins, then freeze — and epochs=5 is phase two, five supervised passes over the frozen encoding; one call runs both, in that order.

Supervised Learning#

Every chapter until now drove the lane with RU, the simplest unsupervised instruction, and paired with the FF read, it gave us unsupervised AHaH plasticity and attractor states that turned out to be logic functions. Supervised learning is more traditional, and in many cases more useful.

Take one labelled flower, AAT encode it, activate all three lanes, and read each with a full-voltage FF to get its activation; then provide conditional feedback depending on the label:

kT-RAM Classifier Routine
aat = encoder.encode(flower)
for lane in range(3): # three labels, three lanes
y = core.evaluate(aat, "FF", lane)
if lane == correct_label:
core.evaluate(aat, "RH", lane) # this IS the class — drive the answer up
elif y > 0:
core.evaluate(aat, "RL", lane) # wrong lane caught saying yes — drive it down
else:
core.evaluate(aat, "RF", lane) # wrong lane correctly saying no — leave it be

The supervised rule has three cases. The lane that owns this flower’s species gets RH, driven up toward yes — these inputs should make it fire. A lane that does not own the flower but said yes anyway is a false positive, so it gets RL and is driven down toward no . A lane that does not own the flower and stayed below zero is already right, so it gets RF, a plain reverse read that undoes part of the anti-Hebbian FF — the kT-RAM equivalent of “all good bro”.

This is essentially a perceptron written in kT-RAM instructions: punish the lane when it is wrong, reinforce it when it is right, one example at a time, with no batch and no stored gradient. The same rule goes by different names in different fields. The neural-net literature calls it the delta rule or the Widrow-Hoff least-mean-squares update, and anyone coming from optimization calls it online or stochastic gradient descent. A neuroscientist would call it an error-driven Hebbian update — fire-together-wire-together, gated by whether the answer was right. They are all the same move: read the output, compare it to the label, and shift the weights a notch in the direction that would have helped.

The winner is the prediction#

Training drives the synapses with full-voltage reads, because you want those reads to move the weights. In this particular routine, we are either rewarding with RH, punishing with RL, or ‘regularizing’ with FF-RF . Inference is the opposite — you want the prediction without disturbing what you learned. So read each lane with FFLV, the sub-threshold read that reports the weight without changing it, then collect the three outputs and take the winner:

from ktram_neural_core.recode import Winner
scores = [core.evaluate(aat, "FFLV", lane) for lane in range(3)]
prediction = Winner().recode(scores) # (argmax,) — the loudest lane wins

Winner is just argmax — the lane with the highest output is the predicted class. Reading three lanes and taking the most active is a collective operation: it happens across lanes, not inside one. Putting it together, the entire classifier is short:

from ktram_neural_core.classify import LinearClassifier
clf = LinearClassifier(encoder, labels=[0, 1, 2], model="byte", init="low", seed=0)
clf.fit(X_tr, y_tr, epochs=5, encoder_epochs=5) # adapt + freeze, then train
pred = clf.predict(some_flower) # encode -> read lanes -> argmax

model="byte" puts an “8-bit memristor” at every synapse, so this is the lane running on a device with real quantization. Swap this out for model="mss" or the other model types to see how it works with different levels of device realism.

Technical abstraction layers#

Everything we have done thus far has driven kT-RAM instructions by hand, one instruction at a time, one lane: read the analog y and branch on it — FF, look at the sign, then apply RH/RL/RF depending on the labels — or use FFLV for non-adapting inference and take the winner. Call this low level L0 — the bare kT-RAM instruction set, one instruction on one lane at a time, the level where you find out what the parts do. Reach for it in a single-synapse lesson, or any experiment where you want to watch every read.

But we will run this supervised classifier routine constantly from here on — it reads a whole group of lanes first, then acts on what they say: to teach, it uses the labels to drive the right lane up and the confident-wrong ones down, and to classify, it takes the winner. It also cheats: the Python reaches into the emulator and reads each lane’s analog output Vy as a floating point number. Hardware could do the same — put an analog-to-digital converter on every lane — but an ADC per neuron is exactly the bad idea I have already harped on at length, and if lane outputs are going to feed anything downstream, they need to come back as AATs, not floats.

We call the hardware that turns a group of analog lane voltages back into AATs an “AAT Recoder” — that covers the reading half of our classifier. The teaching half applies kT-RAM instructions conditionally, branching on the read voltage and the supervised labels. Wrap both halves behind a clean digital interface and you have a second technical abstraction layer, L1: the layer we will actually build in commercial hardware, where the instruction set gives way to fixed, efficient routines. Our first one is RankCut, named for its readout — rank the lane outputs, cut the list to the most active — which, combined with the three-case feedback, is what makes this particular recoder a classifier. It comes as one object with two calls, adapt to teach and read to answer:

from ktram_neural_core.aat_recoder import RankCut
rec = RankCut(core, labels=[0, 1, 2])
rec.adapt(aat, {label}) # teach: FF read on each lane, then RH/RL/RF by label
rec.read(aat) # answer: FFLV read on each lane, recoded to an output AAT

Underneath, those two calls are the exact instructions we wrote by hand a moment ago — same FF, same three-way feedback, same low-voltage read.

In full, the readout returns the addresses of the lanes above a threshold — zero volts, say — strongest first, and cuts the list after at most N entries.

The rank-cut in hardware#

Taking the single winner is the easy readout. Often you want more — the top two or three in order, or every lane that came out positive, ranked. Those are all one operation — sort the lanes by output, drop the ones below a threshold, and stop after at most N — the rank-cut, whose smallest setting is the winner: threshold at the floor, N of one.

Sorting analog values the obvious way means an analog-to-digital converter on every lane and a digital sort algorithm behind them, which would take a pile of silicon — a much cheaper trick exists, and it’s the kind that makes hardware fun (at least for me!).

Every lane hands you a voltage — its output Vy, sitting somewhere between −V and +V. Take one reference voltage, shared by all the lanes, start it above the top of the range, and sweep it down. Each lane has a comparator watching its own voltage against that falling reference, and the instant the reference drops past a lane’s voltage, that comparator trips and the lane calls out its address.

Picture a flood draining off a landscape, the waterline falling until the highest peak breaks the surface, then the next, then the next; write down the order they appear and you have sorted them tallest-first without measuring a single height. The swept reference is the waterline, the lanes are the peaks, and the strongest lane trips first, the next strongest second, with the addresses arriving in rank order. The sweep stops when the reference reaches the threshold, since everything still underwater is a no, or after N lanes have surfaced.

A two-panel comparison. On the left, five lanes each feed an identical multi-bit ADC block containing a comparator, a capacitor bank, and conversion logic; each block emits an 8-bit binary number, and all five numbers flow into a digital sort block that outputs sorted digital outputs. On the right, the same five lanes are vertical bars of shuffled heights, each topped by a single comparator. One dashed waterline crosses all five: lanes 3, 1, and 4 rise above it and carry rank badges 1st, 2nd, and 3rd, while lanes 2 and 5 sit below it, untripped. An arrow labels the waterline as the shared reference sweeping down toward a threshold where the sweep stops, and the panel's output arrow reads digital addresses out in sorted order.
Two ways to digitize a bank of lane outputs. The ADC route on the left gives every lane its own converter and ships the resulting numbers to a digital sort engine. The rank-cut on the right shares one falling reference across the whole bank: each lane keeps a single comparator, the lanes trip in descending order as the reference passes them, and the addresses come out sorted. Lanes still below the waterline when the sweep stops are never reported.

The ADC route puts a full converter on every lane — a comparator, a capacitor array, and conversion logic — then ships all those digitized numbers off to a sort engine with the memory to hold them. That shipping is the expensive part: most of the energy in CMOS goes into charging and discharging the wires that carry bits between blocks, not into the logic itself. The swept reference keeps everything local: one ramp shared across the array, one bare comparator per lane, and a latch to record the order. The sort is done by the time the ramp finishes, encoded in when each lane trips rather than in a number you had to compute and move, and since the answer is the firing order, you can stop the ramp the moment you have your N or hit the threshold — an easy decision settles early. One sweep of time saves a lot of silicon and a lot of energy, exactly the trade you want on hardware meant to save energy.

Does it work?#

Accuracy on its own says little; the relevant measure is how the lane compares to a real linear classifier given the exact same inputs. The AAT encoder is doing some work and a fair test has to hold it fixed. So I freeze the A2D encoder once and run three things on the identical AAT encoding: our kT-RAM lane, scikit-learn’s LogisticRegression, and a LinearSVC. All three are linear classifiers. The reference two solve for their weights in one batch over the whole dataset. The lane learns online, one example at a time, with local instructions. A fourth bar, plain logistic regression on the raw measurements, shows what the encoding itself costs or gains us. The question is whether the new L1 RankCut routine lands where the batch solvers do.

A bar chart of test accuracy across twenty seeds. Four bars: LogReg on raw features at 0.963, set apart on the left; then on the same AAT encoding, LogReg at 0.961, LinearSVC at 0.951, and the kT-RAM lane at 0.962. The three encoded bars sit at essentially the same height, with overlapping error bars.
The kT-RAM lane against two batch linear classifiers on the identical AAT encoding, averaged over twenty train/test splits. The lane lands at 0.962, right on top of logistic regression's 0.961 and a hair above the linear SVM — an online, local, instruction-level rule matching batch solvers that see the whole dataset at once. The raw-feature bar on the left shows the encoding barely moves a linear classifier's accuracy here, neither helping nor hurting much on this data.

Each lane sees only its own weights and its own read, yet that online rule still lands on the same accuracy as a batch solver holding the entire dataset in memory.

Two confusion matrices side by side for one train/test split, the kT-RAM lane on the left and logistic regression on the right. Both are identical: a clean diagonal of 13 setosa, 13 versicolor, 12 virginica, every off-diagonal cell zero, accuracy 1.000.
One split, the kT-RAM lane and logistic regression side by side on the same encoding: both are perfect here, every setosa, versicolor, and virginica on the diagonal and nothing off it — but not every split is clean, so a single perfect split can flatter. The twenty-seed average above is the number that counts, and the next figure shows where the misses land.

Where the misses land#

The twenty-seed average was 0.962, not 1.0, so on some splits the lane misses a flower or two — which flowers, and why? Iris has four measurements and the lane reads all four at once, so any 2D plot is one shadow of that four-dimensional data; here are three of them:

Three scatter plots of the same Iris flowers, each a different pair of the four measurements: petal length vs petal width, sepal length vs sepal width, and petal length vs sepal length. Setosa, in blue, sits in its own clump in every panel. Versicolor in green and virginica in red overlap heavily — one diagonal band along the petal axes, fully intermingled in the sepal panel. Two test flowers the lane got wrong are circled in black, and in every panel they sit right where green and red meet.
The same four-dimensional Iris data, three pairs of measurements at a time. Setosa is linearly separable in every view, so no classifier ever misses it. Versicolor and virginica stay tangled in all of them, worst in the sepals and cleanest along the petals. The two flowers the lane missed are circled, and they sit on that green/red boundary in every panel. No straight cut splits them. A small error is the best any linear classifier does here, and the lane matches logistic regression and the linear SVM.

The lane and logistic regression miss the same flowers, the ones in the versicolor–virginica overlap — the support vectors from Chapter 5b that no single line can satisfy at once, a property of the data rather than the learning rule.

The ceiling is expected: one cut per class, learned with local instructions, lands exactly where the standard linear solvers do. Getting past that straight line takes more than one cut, and we will get there soon.

Reading at temperature#

Inference read each lane with FFLV — quiet enough that the winner comes out the same on every read — but every kT-RAM read carries the kT-bit’s read noise, and we have two dials for it : the read voltage and the read pulse width. The thermal part of the hiss grows as 1/V1/V, and it grows again as the pulse gets shorter, because a longer read integrates the hiss down. We turn the voltage here, with a noise fraction that scales it down by 1−noise1-\text{noise}: at the standard low-voltage read the result is nearly clean, and as VV drops toward zero the hiss swallows the signal.

y = core.evaluate(aat, "FFLV", lane, noise=0.6) # read at (1 - 0.6) x 0.05 V = 0.02 V

Turn that dial and the winner is no longer settled in advance: each lane reports its true voltage plus a random kick, and the argmax becomes a draw. The draw favors the loudest lane, but a close runner-up with more noise can take the win. For a setosa nothing changes: lane 0 is so far ahead that no plausible kick unseats it. For a flower on the versicolor/virginica overlap, the two lanes read within a whisker of each other, and the hot read returns versicolor on some reads and virginica on others. The classifier no longer returns a verdict; it returns a sample.

Whether those samples mean anything depends on what the weights hold, and one line of the training routine decides that.

Soft and hard feedback#

Go back to the three-case rule and look at the middle case: a wrong lane that fired anyway gets RL, driven down. That one line decides what the lanes become.

hard feedback
for lane in range(3):
y = core.evaluate(aat, "FF", lane)
if lane == correct_label:
core.evaluate(aat, "RH", lane)
elif y > 0:
core.evaluate(aat, "RL", lane)
else:
core.evaluate(aat, "RF", lane)
soft feedback
for lane in range(3):
y = core.evaluate(aat, "FF", lane)
if lane == correct_label:
core.evaluate(aat, "RH", lane)
else:
core.evaluate(aat, "RF", lane)

Keep it and the feedback is hard: a lane is punished every time it speaks out of turn, so the stable outcome is one loud lane per region of input space and every other lane pushed below zero. Hard feedback carves the decision boundary, maximizing the decision margin.

Drop it and the feedback is soft: the wrong-but-fired lane falls through to the RF case with everyone else, so no lane is ever driven down for firing. A lane climbs when its label is the right answer and holds its ground otherwise, so where the classes genuinely overlap, both lanes stay above the threshold instead of one being beaten under. Hard feedback keeps only the single best answer and discards the rest; soft feedback retains every answer the data supports. In the emulator this is one argument: rec.adapt(aat, {label}, feedback="soft").

Hard weights and a cold read are the typical classifier’s domain, where we want to maximize the decision boundary and punish indecision, while soft weights and a hot read are a different machine: shown a flower from the overlap, it answers versicolor on some reads and virginica on the rest, because the training data never settled the question either. Because the noise is conditioned by the weights, the samples carry more variance around the areas of uncertainty.

Label in, pattern out#

A group of kT-RAM neural lanes runs one direction: an AAT goes in, and an AAT comes out, and nothing says which side the labels sit on. The classifier learned the pattern-to-label mapping, so you cannot hand those trained lanes a label and get a flower back — the mapping is the wrong direction. A generator is a new set of lanes that learns the opposite mapping, trained separately, on the same data, with the same basic supervised routine. The label enters as a plain input coordinate, and the output channels stand for pattern bins.

Take a flat two-dimensional dataset — points on a plane, each carrying a class label. Bin both axes with the fixed-bin A2D from earlier, five bits per axis for thirty-two bins each — the widget below has a resolution knob, so you can sweep it. To classify we did the obvious thing: one lane per label, reading the bins. To generate, swap the roles: thirty-two x lanes, one per x bin, read only the label, and thirty-two y lanes, one per y bin, read the label together with the x bin — together those sixty-four lanes are the generator. Teach them soft, and teach them with slightly hot reads too: the hiss dithers the byte-quantized updates, smearing hard rounding thresholds into smooth averages, the same dither trick audio engineers have used for decades, done here by the physics itself.

The classifier’s lanes start at the low init, every conductance near its floor; the generator lanes start at the medium init, every conductance near mid-scale. A lane’s hiss scales as one over the square root of the total conductance the read sees, so the smallest-magnitude lane is the loudest — start the generator lanes at the floor and every untrained lane out-shouts the trained ones on every hot read, and the samples are noise. The mid-scale start gives every lane a starting conductance instead, so an untrained lane reads near zero and stays quiet until training moves it.

To draw one sample, clamp a label and read the x lanes hot: the most active lane is an x bin — commit it. Then read the y lanes hot, given the label and that committed x, and the winner is a y bin; decode the two bins back into a point, and that is the entire generator.

The order of reads matters: the number of steps equals the number of tuples in the AAT we are generating. This dataset is two-dimensional, so the output AAT has two tuples, an x symbol and a y symbol, and one sample takes two reads — an AAT with kk tuples takes kk reads. The rule at every step is the same: read the lanes for one tuple given the label and every symbol committed so far, take the most active, commit it, and move to the next tuple. That is the chain rule of probability, p(x,y)=p(x) p(y∣x)p(x, y) = p(x)\,p(y \mid x), running as lane reads, and from here on we call the whole sequence of reads the chain.

The split into x lanes and y lanes is how the chain is organized, not a hardware requirement: a lane simply ignores any input space held at NONE, so you can pool everything into one array of lanes whose input spaces cover the label, the x bin, and the y bin, with one lane per symbol across all three — sixty-seven lanes here. The chain runs the same way: clamp what you know, hold the rest at NONE, read the lanes of one unknown space, commit the most active into that space, and move on to the next. That is an auto-associative array — clamp any subset of the entries and the chain fills in the rest, and clamping the label draws a pattern. The generator in the widget is this array with its rows split apart, which keeps the display legible.

The sequence captures the agreement between the tuples. Lanes that see only the label can report one thing — a marginal: how often each of their bins gets used by that class. They have no way to report that y runs high whenever x runs high, because nothing in their input says which x we are on. Committing x changes the question the y lanes are asked: instead of “where does this class put its y mass,” it answers “where does this class put its y mass on the slice of the plane at this particular x.” If the class is a tilted cloud, those two questions have different answers, and only the second one draws the tilt. Skip the commit and read the x and y lanes at the same instant, each side seeing only the label, and every draw comes from the two marginals instead — the histograms still match, but the tilt flattens into an axis-aligned blur, because nothing ties any particular y to the x it came with. The widget below has a switch for exactly this, chained against synchronous, so you can watch the correlation appear and vanish. Run the same chain over hundreds of tuples instead of two and you have the image generator a few chapters from now, each read filling in one more piece of the picture while looking at every piece already filled in.

Press Start, and lanes train and generate 2D Gaussian clusters live in your browser. The left panel holds the data and the classifier: you can add, drag, tilt, and relabel the Gaussian clouds, and three label lanes trained hard underneath color every cell of the plane by its cold winner. Drag a cloud across the plane and the boundary chases it, because the lanes never stop learning. The right panel is the generator — the soft x and y lanes drawing samples at whatever temperature you set, with marginal histograms along the top and side comparing what it draws against what the data holds.

kT-RAM Neural Lane Classifier-Generator Demo
stopped — press Start to train the lanes live
CLASSIFIER — hard feedback · cold read
click to add a cloud · drag to move · curves = data marginals
GENERATOR — soft feedback · hot read
bars = generated marginals · grey = the data's
mixture:resolution:bits
click a cloud to select it · click empty ground to add one
draw:
temp0.15
byte-model lane arithmetic — TwoOne divider, FF/RH/RL/RF updates, read-noise sampling
Byte-model kT-RAM lanes learning a mixture of Gaussians, live in your browser. Both panels run the same integer lane arithmetic in opposite directions: the left reads a pattern to name a class, the right reads a class to draw a pattern.

Try a few things: drag the temperature to zero and the samples collapse to one spot per class, since with no hiss the chain can only ever pick its single loudest bin pair; raise it and the cloud fills back out, and push it far and the samples start spilling past the data. In between, the histograms settle onto the grey reference curves, and warmer reads trace the shape better than cool ones, because the noise is doing the mixing. Load the tilted preset and the generated cloud comes out visibly tilted, the chain carrying the correlation. Now flip the draw from chained to synchronous: the x and y lanes read at the same instant, each side seeing only the label, and the tilt collapses into an axis-aligned blur while the marginal histograms barely move. The joint structure was never held in the x lanes or the y lanes alone — conditioning each y read on the committed x is what preserves it.

Then flip the feedback to hard and watch what the punishment does: the RL case starts beating down every lane that fires out of turn, and in a generator that is nearly every lane on nearly every example, because the answer is genuinely spread across many bins. The read always returns its most active lane, so the samples keep coming, but within seconds they stop looking like the data — a generator punished for every wrong guess cannot hold a spread, only a single best answer, with everything else beaten flat. Flip back to soft and the distribution fills back out.

Where we are#

We used three lanes over an A2D encoding, one lane per label, trained with a simple kT-RAM instruction set routine: read cold, that lane matches batch logistic regression and a linear SVM, and read hot, the same routine keeps the odds instead of the verdict. Teach fresh lanes the opposite mapping, chain the reads, and they draw patterns instead of naming them.

Two memristors wired against each other and a handful of voltage patterns have done a lot so far: they held a bit against the thermal bath in Chapter 3, were memory or logic or inference by choice of partition in Chapter 4, assembled logic gates from the structure of the data in Chapter 5, and found the maximum-margin boundary in Chapter 5b. Here they land on logistic regression, then sample a joint distribution when softened, read hot and sequenced, all with no new devices, circuits, or instructions added.


Next: Chapter 6b: The Unsupervised Basis Encoder