kt-ram · ai-hardware · assimilaate · aat-codec · vector-quantization · transformers
Chapter 7: The AAT Codec
kT-RAM processes AATs while modern AI models process real-valued vectors, so we need a fast, accurate translator between them. This chapter builds the first AAT Codec.
By Alex Nugent ·
Contents
The iris flowers made encoding easy: four measurements, four rulers, each measurement dropped into a bin on its own ruler, and four numbers became a four-symbol AAT.
It worked because the four measurements are separate, and a wide petal tells you nothing about a narrow sepal, so binning each one alone throws almost nothing away.
Most vectors are not like this, because a transformer’s vectors have dimensions that move together.
So how do you turn a whole vector into an AAT?
Binning a whole vector#
Start with two numbers so we can draw it: a vector in , one point on a plane. When is large, usually is too, so the points form a diagonal streak.
Bin each axis: eight slices across, eight up, sixty-four cells, and the data only visits the cells along the diagonal. You laid out sixty-four cells and used twenty-three, because the two rulers are set independently.
The A2D encoder from the last chapter softened this by sliding each axis’s bin edges toward where the data is densest, which is enough for four loosely related measurements. But an adaptive ruler fixes the spacing within an axis, and the diagonal is a fact about two axes at once.
One address for a whole vector?#
The data clumps, so put the addresses on the clumps and drop a few dozen reference points, called centroids, where the vectors sit. To encode a vector, find its nearest centroid and write down that centroid’s index: “nearest to centroid #37” instead of “bin 3 on , bin 5 on .”
K-means finds those clumps: give it vectors and ask for S clusters, and it places S centroids so every vector ends up near one. Replacing a vector with its nearest centroid is vector quantization, which is decades old and runs inside everything from early speech coders to modern audio, and every field that reinvented it renamed the parts. Compression calls the set of centroids a codebook and each entry a codeword, sparse coding calls it a dictionary of atoms, and all of them mean a learned, indexed list of representative vectors. We say centroid, and codebook for the list that holds them.
This does not scale: to address as finely as we addressed that square, the number of centroids explodes, because high-dimensional space is nearly all corners and emptiness, the curse of dimensionality. A centroid near every vector would need a codebook bigger than the model itself, so one symbol cannot carry a whole 128-number vector.
Cut it into slots#
So cut the vector into pieces and address each piece. Chop into 64 slots of two numbers each, so every slot is a small vector in , the space we just learned to handle. Give each slot its own codebook of 64 two-number centroids, fit with k-means on that slot’s data. Then encode each slot on its own: slot 0 is nearest its centroid #12, slot 1 its centroid #50, and so on, making the vector 64 integers, an AAT.
Each slot’s centroids capture the correlation two dimensions at a time, and the codebooks stay small: 64 centroids per slot, not one enormous book for the whole space.
Data-science readers know this as product quantization, which drives billion-scale similarity search and is the part of FAISS that finds near-neighbors in a corpus too big to hold in memory. A vector in becomes k = D/2 slots, and each slot carries one symbol out of S = 2^bits, the index of its nearest centroid in that slot’s codebook, so the whole code is k symbols, or k · bits bits on the wire. Write for one slot’s alphabet, the symbols . An AAT is a point in : slots, choices in each, the same way a real vector is a point in . The notation reads like coding theory’s , where the subscript is the alphabet size and the superscript is how many symbols you string together.
For 64 slots of 6 bits each, the AAT is bits, or 48 bytes: a point in . The vector it came from takes 256 bytes in 16-bit floats, five times as much. A saved byte is never stored, never crosses a memory bus, and never draws energy.
kT-RAM accepts AATs and nothing else, so all computing on this hardware happens in AAT space, , while models (at this time) are trained and shipped in . We had better be good at going back and forth:
An encode–decode pair is a codec, the same word and much the same job as the audio and video codecs that pack a symphony into an MP3 and get it back out. This chapter builds an AAT Codec library and tool set, and the product-quantization codec we just assembled is Codec A.
Codecs out of and into address space are all over machine learning now, where the FSQ and BSQ image tokenizers round onto uniform rulers or take the sign of each axis. The neural audio codecs behind modern voice models give every few milliseconds of sound a small stack of addresses.
A codec has two jobs, and the first is the round trip: encode to , decode back to , lose as little as possible. The second is kernels, mathematical operations run directly on the codes in , with no decode at all. Score two vectors, sum them, take a norm, all in symbol space. These kernels take fewer elementary operations than the float versions they replace, touch narrower numbers, and read from tables that fit in hundreds of bytes. We convert to AATs so kT-RAM can assimilate transforms. The codes are also worth computing on directly.
Reconstruction Loss#
Take a vector from the model, encode it, then decode it straight back to : look up the centroid at each slot’s address and lay them end to end, then measure how close the result is to the original.
We ran Codec A on real activations from a working transformer, Qwen3-0.6B, taking one middle layer with codebooks fit on one split of the data and scored on another it never saw. The measure is cosine similarity, so at 1.0 the decoded vector points exactly where the original did, and lower means it drifted.
At six bits we land near 0.98 cosine, most of the fidelity at a fraction of the size. One caveat: cosine only checks direction, so a code can get the direction right and the length wrong. Parts of a transformer depend on length, so the next chapter measures both.
Whole images to AATs#
Codec A works on any vector, images included. The 28×28 Fashion garments from the last chapter are 784-number vectors, so a garment can be cut into slots and addressed like anything else. Chapter 6b already did a version of this, cutting each garment into sixteen 7×7 tiles and giving every tile the address of whichever learned feature won that tile’s read, so a whole garment came back as sixteen symbols, and slots and tiles are the same move at two scales.
Let’s try another way: instead of breaking the image into pieces, encode it whole, keeping the shape of , books with entries each, and changing what an entry represents. An entry is now a full 784-number image, a faint ghost of a whole garment rather than a piece. It is one term in a sum, nothing clusters around it, and alone it looks nothing like a garment. Sparse coding calls it an atom, the code holds one address per book, and the decode adds the selected atoms. Every symbol now influences every pixel.
Fitting the books is k-means one level up, where k-means assigns each point to its nearest centroid, then moves that centroid to the mean of its assigned points, and repeats. Here the assign step encodes every training garment, so each one ends up with addresses, while the move step changes every atom to whatever makes the sums it appears in come out closest, and the two repeat. The one difference is that a garment is rebuilt from a sum of atoms rather than a single centroid, making the move a joint solve instead of an average. Two atoms that keep turning up in the same garments’ codes must divide the work.
The move step makes an atom look like a garment, because an atom is a term in the sum of every garment whose code holds its address, so the solve drives it toward the structure those garments share. Anything else it carried would throw all of them off at once and raise the error being minimized, and so the early books carry the silhouette of the whole garment while the later ones carry the detail left over. The figure below uses books of atoms, so a garment becomes a point in , or 512 bits, 64 bytes.
This is the additive family of codes. Babenko and Lempitsky named additive quantization in 2014, the two-step fit is the loop dictionary learning has always used, and the family now carries frontier LLM weight compression, recommender item addresses, and production speech codecs. The next chapter takes up what these codes cost, how far the fit can be pushed, and how they measure against the patch route. The gray-scale garment reconstructions ending Chapter 6b came from a network using a two-level version of this additive AAT encoding.
Staying in AAT space#
Both constructions map between and , but for the transformer the round trip is only the start of the problem. Every encode and decode costs time and energy, so once we encode we stay in AAT space as long as we can, decoding only when an operation or the end result needs real numbers.
kT-RAM consumes and produces AATs, so to assimilate a transformer we need a plan for the operations that sit outside the layers a neural lane can represent:
- inner products between vectors — the attention score
- vector length — the norm, used in normalization and scoring
- weighted sums of vectors — the attention output, blending many value vectors into one
- addition — the residual stream running up the spine of the network
None of these operations use the learned parameters, but a transformer cannot work without them. The projections that make the queries, keys, and values are big matrix multiplies, and so is the MLP between attention blocks. The kT-RAM neural lanes can assimilate those learned parameters, while these four other operations run between the lanes, and the longer the context, the larger their share of the arithmetic.
That leaves two routes: decode back to and re-encode later, or run these operations directly in AAT space. The second route is better.
Take the attention score, the inner product of a query vector and a key vector: multiply the two element by element and sum the products. The obvious route decodes both vectors back to and performs 128 multiply-accumulates. An AAT codec kernel can do the same job with no multiplies at all: precompute the inner product of every centroid pair, slot by slot, and store them in lookup tables, and scoring two codes is then 64 table reads and 63 additions, with vectors never decoded. In gates and energy, a multiply-accumulate is among the most expensive operations in digital hardware, and a table read and an add among the cheapest.
Codec A’s costs#
Codec A’s cost shows up in two places. First, the lookup tables are large, because scoring queries against keys needs a table of inner products for every slot, , over a quarter-million entries for one layer’s score kernel. Second, encoding is worse, because finding each slot’s nearest centroid means measuring the distance to all 64 of them, so encoding one vector costs thousands of multiplies. Getting into takes far more arithmetic than the 128 multiplies the kernel saved!
Codec A works but it is unshippable. Can an AAT codec keep this fidelity with tables small enough to sit next to the hardware and an encoder cheap enough to run on every vector coming down the residual stream? That is the next chapter.