Loading the catalog…
Loading the catalog…
Reducing Hit Time Hit time is the time to find data that is in the cache: use the index to pick a set, read its tag(s) and data, compare tags, then choose the right block and send it to the CPU. Both techniques here make that path shorter. 1. Small and Simple Caches Small helps because a smaller memory array has shorter wires and fewer rows to decode, so indexing and reading are faster. This is why L1 size often stays flat across processor generations (as in the AMD K6, Athlon, Opteron example). Making L1 bigger would slow every access, and hit time sets the clock cycle. Simple (direct-mapped) helps because each address has only one possible location. The cache can send the data to the CPU right away, while the tag check runs in parallel. If the tag turns out not to match, the data is thrown away. A set-associative cache can't do this. It must finish all its tag comparisons before the multiplexer knows which way's data to send. 2. Way Prediction The problem it solves There are two options, and each has a weakness: Direct-mapped: fast hit, because there is one possible location and the data can be sent before the tag check finishes. But it has more conflict misses. 4-way set-associative: fewer conflict misses, but slower hit, because it must compare 4 tags and then use the mux to choose the data. Way prediction keeps the set-associative cache, and its low miss rate, but guesses which way holds the data. If the guess is right, the access behaves like a direct-mapped one. "Use load PC or XOR the load src reg and load offset to index the prediction table" The prediction table is a small memory that remembers which way each load found its data in last time. It has to give an answer before the cache access starts, but the real memory address (base register + offset) isn't finished computing yet. So the table is indexed by something available earlier: the PC of the load instruction. The same load instruction tends to hit the same way again, for example inside a loop. or base register XOR offset, a quick stand-in for the address. XOR has no carry chain, so it's much faster than the full addition. correct prediction Wrong way, hit elsewhere Miss in all ways Random access(85% accurate) With 85% accuracy, most hits run at direct-mapped speed. You can estimate the average hit time like this: Average hit time ≈ 0.85 × 1 cycle + 0.15 × 2 cycles = 1.15 cycles The miss rate stays that of a 4-way cache, because every way is still checked before declaring a miss. Only the order of checking changes. "Any other benefit of way prediction?" — power savings If you click random a few times, the average tag compares per access drops far below 4. A normal 4-way cache powers up all 4 tag comparators and reads all 4 data blocks on every access, then throws away 3 of them. With way prediction, a correct guess needs only 1 tag compare and 1 data array read. That is roughly a quarter of the dynamic energy per access in a 4-way cache. The saving matters a lot in L1 caches, which are accessed almost every cycle, and in battery-powered chips. Some designs go further with way selection: they read the data array of only the predicted way. Example just read, not in the exame (a) For a 1024 KB L2 cache with 64-byte blocks and 8-way set associativity, how many way prediction table entries are needed? 1024 KB, 64-byte blocks, 8-way Blocks = 1024 KB ÷ 64 B = 220 ÷ 26 = 214 = 16,384 blocks Sets = 16,384 ÷ 8 = 211 = 2,048 sets Bits per entry = log2 8 = 3 bits Answer: 2,048 entries, each 3 bits, for a total of 2,048 × 3 = 6,144 bits (6 Kb). (b) For an 8 MB L2 cache with 128-byte blocks and 2-way set associativity, how many way prediction table entries are needed? (b) 8 MB, 128-byte blocks, 2-way Blocks = 8 MB ÷ 128 B = 223 ÷ 27 = 216 = 65,536 blocks Sets = 65,536 ÷ 2 = 215 = 32,768 sets Bits per entry = log2 2 = 1 bit Answer: 32,768 entries, each 1 bit, for a total of 32,768 bits (32 Kb). (c) What is the difference in the way that the processor with only 8Kb way prediction table will support the cache in part (a) versus the cache in part (b)? Cache (a) is fully supported. 6,144 bits fits in 8,192, so every one of the 2,048 sets gets its own 3-bit predictor, with about 2 Kb left unused. Prediction accuracy is as good as the predictor allows. Cache (b) is only partially supported. The table can hold only 8,192 one-bit entries, but there are 32,768 sets. The table is indexed with only the low 13 bits of the 15-bit set index, so 4 different sets share each entry (aliasing). When those sets are used in the same period and their blocks sit in different ways, they overwrite each other's prediction. This causes more mispredictions, and each one costs an extra cycle to check the other way, so the average hit time rises. The cache still works correctly in (b). A prediction is only a hint, and the tag comparison always checks it. A wrong guess costs time, not correctness. If your course means "8K entries" rather than 8K bits, the conclusion is the same. Cache (a) needs only 2,048 of the 8,192 entries (with 3-bit entries), so it fits. Cache (b) needs 32,768 entries, so each entry is still shared by 4 sets.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
Reducing Hit times. Reducing Hit Time Hit time is the time to find data that is in the cache: use the index to pick a set, read its tag(s) and data, compare tags, then choose the right block and send it to the CPU. Both techniques here make that path shorter. 1. Small and Simple Caches Small helps because a smaller memory array has shorter wires and fewer rows to decode, so indexing and reading…
Open source