There’s definitely strategies here; A lot of the floating point operations use subnormals, and a lot of the worst instructions are slowed down by really, really fucking with MMIO.
This author also has other things like: A compiler that emits only `mov` instructions and another compiler that deliberately messes with the control flow so that, if disassembled, common debuggers will draw symbols like skulls or threats. https://github.com/xoreaxeaxeax/repsych
Also came up with the original ..cantor.dust.. binary visualization tool, which is a tool I never used directly but Chris' presentation of it in 2012 is still one of the coolest talks I've ever seen.
It is literally impossible to respond to input in 10ms on most platforms, for various reasons. The USB input lag of 12-30ms and the 60Hz refresh rate of most monitors being just the first two.
60Hz monitors definitely prevent it, but I'm pretty sure USB lag is far less than 12-30ms. My USB mouse can make a round-trip to a remote server faster than that.
The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]:
- 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.
- 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data.
- 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect.
That’s the prevailing statistic, but your original number isn’t that wrong either:
Humans can perceive much smaller latencies.
If you look at the Card & Miller reference, at least some humans can perceive differences in ~50ms vs 100ms latencies when typing (in my limited testing, it’s likely you can!). There’s some newer research I don’t have handy that I believe found error rates decreased and NSAT improved until around at least 30ms (if not 20ms).
On that note, humans can definitely distinguish 60hz vs 120hz reliably (about 8ms faster per frame).
Even faster: with a reference (eg when dragging on a touchscreen), humans can distinguish down to at least 1ms vs 10ms of latency: https://m.youtube.com/watch?v=vOvQCPLkPt4
And you can probably distinguish metronomes that are off by about 1-2ms. Much smaller for other things (like metronomes that slightly slower or faster than one another).
Sorry for nitpicking but the logical inverse of this statement is the stronger and more appropriate version:
If some app does not respond in 10ms or less, it is not interactive.
The new mspaint fucked, then unfucked, then refucked my decades-old muscle memory of Win+R mspaint Enter Ctrl+E 1 Tab 1 Enter Ctrl+V to open Paint, resize canvas to minimum, then paste from clipboard. When you press Ctrl+E now, the Units control is selected by default, for some completely asinine reason!!!
There are two aspects to compute performance: latency and throughput.
There has been a ton of work to improve throughput, but that's often come at the cost of latency, because a great technique to improve throughput is batching, to avoid the per-task overhead cost.
We've also added many layers of abstraction.
A seminal example is Dan Luu's "computer latency" table (2017) https://danluu.com/input-lag/, which measures the time between a keypress and visible changes on the screen, and wherein an Apple IIe has latency 5x lower than a Lenovo X1 Carbon on Windows.
From some perspective, you could say that the Lenovo on Windows example is an impressive feat of engineering, considering all the subsystems involved (USB, interrupt dispatcher, input layer, windowing system, double-buffering...)
This throughput-over-latency tradeoff can be seen across the entire computer landscape design.
It'd be really interesting to see whether the winning (losing?) instructions/strategies would be different on other architectures. At least right now the top spot (`fxrstor64` on MMIO, starve PCIe) seems relatively architecture-independent, but maybe something about MMIO ordering rules on e.g. POWER would be different enough to change that -- or perhaps open up new avenues?
I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.
It's not an implementation detail, because the decoder runs before the execution of every instruction. If we're going to say that NOP increments IP by one, then we should also say that ADD "stores in dst the addition of src and dst, as well as incrementing IP by the length of the instruction", and JMP imm "increments JMP by imm + the length of the instruction".
I'm not disputing the total effect. I'm asking if you'd rather describe ADD and JMP in this manner, in order to say that NOP does not in fact do nothing.
As specified by the spec, it arguably increments RIP by one.
The actual typical hardware implementation just fetches the next 16-32 bytes from icache, shifts it to the correct alignment, and slams it into a bunch of parallel decoders which each attempts to decode one x86 instruction per byte.
The next cycle, the first 1-6 non-overlapping valid instructions are accepted into a queue for further decoding. The NOP almost certainly takes up space in this queue.
At no point does RIP get incremented by one. There isn't even a single physical RIP register to increment, the CPU is "executing" dozens or even hundreds of RIPs in parallel.
I thought I remembered reading somewhere re: the 8086 microcode disassembly that NOP, which is encoded as XCHG AX,AX actually does run the XCHG microcode and uses an internal scratchpad register to do the exchange.
Chris usually takes an educational angle, I don't think this is LLM-generated content at all it's just his style. I highly encourage watching some of his DefCon or BlackHat talks, they're fun!
Very cool! Also, huh interesting. I’ve used rdtsc to measure cycle diffs but had no idea its execution takes that long. Is that common across architectures?
submitted by two different people, both with year+ old accounts and decent karma. i dont think either is the author. the other submitter probably read this one, looked at the github, saw something else cool and posted it. (i almost did the same, but bookmarked it instead)
Bus cycles can be arbitrarily long on any processor that has memory cycles with a hand shake requiring an ack, with no timeout.
E.g. we can build a board around a MC68000 where we make it lock up forever in a bus cycle, waiting for a DTACK that doesn't arrive.
Some early microprocessors had clocked bus cycles without handshaking. They would put out an address on some address lines and signal some line together with a read/write indication, and then expect the transfer to be completed within some clock cycles. If nothing is attached to the address, they would read whatever values are on the bus, like maybe all 1's if it is an open drain system that requires the transmitting device to pull to ground to indicate zero.
I'd say that kind of thing belongs to a hall of shame; it requires software hacks to interface with anything that can't keep up with the prescribed bus cycle.
Ultimately... failure modes must arise because naive gate design simulation models can't determine issues in a computationally feasible time frame.
The consequences of "fixing" CDC prone design flaws makes a processor many times slower (8 to 16 times slower on my dumb attempt), and develops weird alien design features very different from Von Neumann architectures.
vardump | 10 hours ago
arn3n | 10 hours ago
TomatoCo | 10 hours ago
inigyou | 9 hours ago
vanderZwan | 30 minutes ago
[0] https://github.com/Battelle/cantordust
[1] https://www.youtube.com/watch?v=4bM3Gut1hIk
metadat | 10 hours ago
What’s that law called about programmers wasting all the compute on abstraction?
LoganDark | 10 hours ago
m463 | 9 hours ago
If some app responds in 10ms or less, it is INTERACTIVE.
makes you think.
Xirdus | 8 hours ago
xboxnolifes | 8 hours ago
m463 | 7 hours ago
I looked it up and it is .1 seconds (100ms)
The basic advice regarding response times has been about the same for thirty years [Miller 1968; Card et al. 1991]:
- 0.1 second is about the limit for having the user feel that the system is reacting instantaneously, meaning that no special feedback is necessary except to display the result.
- 1.0 second is about the limit for the user's flow of thought to stay uninterrupted, even though the user will notice the delay. Normally, no special feedback is necessary during delays of more than 0.1 but less than 1.0 second, but the user does lose the feeling of operating directly on the data.
- 10 seconds is about the limit for keeping the user's attention focused on the dialogue. For longer delays, users will want to perform other tasks while waiting for the computer to finish, so they should be given feedback indicating when the computer expects to be done. Feedback during the delay is especially important if the response time is likely to be highly variable, since users will then not know what to expect.
from Jakob Nielsen:
https://www.nngroup.com/articles/response-times-3-important-...
less readable but the original paper:
https://www.yusufarslan.net/sites/yusufarslan.net/files/uplo...
citelao | 6 hours ago
Humans can perceive much smaller latencies.
If you look at the Card & Miller reference, at least some humans can perceive differences in ~50ms vs 100ms latencies when typing (in my limited testing, it’s likely you can!). There’s some newer research I don’t have handy that I believe found error rates decreased and NSAT improved until around at least 30ms (if not 20ms).
On that note, humans can definitely distinguish 60hz vs 120hz reliably (about 8ms faster per frame).
Even faster: with a reference (eg when dragging on a touchscreen), humans can distinguish down to at least 1ms vs 10ms of latency: https://m.youtube.com/watch?v=vOvQCPLkPt4
And you can probably distinguish metronomes that are off by about 1-2ms. Much smaller for other things (like metronomes that slightly slower or faster than one another).
This is a special interest of mine XD
sitzkrieg | 6 hours ago
faresahmed | 6 hours ago
HappyPanacea | 10 hours ago
adamrezich | 8 hours ago
fluoridation | 6 hours ago
sitzkrieg | 6 hours ago
EvanAnderson | 5 hours ago
mwigdahl | 9 hours ago
inigyou | 9 hours ago
summarybot | 9 hours ago
RiverCrochet | 6 hours ago
inigyou | 9 hours ago
AceJohnny2 | 5 hours ago
There has been a ton of work to improve throughput, but that's often come at the cost of latency, because a great technique to improve throughput is batching, to avoid the per-task overhead cost.
We've also added many layers of abstraction.
A seminal example is Dan Luu's "computer latency" table (2017) https://danluu.com/input-lag/, which measures the time between a keypress and visible changes on the screen, and wherein an Apple IIe has latency 5x lower than a Lenovo X1 Carbon on Windows.
From some perspective, you could say that the Lenovo on Windows example is an impressive feat of engineering, considering all the subsystems involved (USB, interrupt dispatcher, input layer, windowing system, double-buffering...)
This throughput-over-latency tradeoff can be seen across the entire computer landscape design.
loeg | 4 hours ago
metadat | 4 hours ago
loeg | 3 hours ago
codeshaunted | 10 hours ago
bee_rider | 10 hours ago
inigyou | 9 hours ago
achierius | 10 hours ago
I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.
spoocecow | 9 hours ago
layer8 | 9 hours ago
jooops1 | 9 hours ago
fluoridation | 6 hours ago
russdill | 6 hours ago
fluoridation | 4 hours ago
loeg | 4 hours ago
fluoridation | 4 hours ago
loeg | 2 hours ago
fluoridation | 2 hours ago
phire | 21 minutes ago
The actual typical hardware implementation just fetches the next 16-32 bytes from icache, shifts it to the correct alignment, and slams it into a bunch of parallel decoders which each attempts to decode one x86 instruction per byte.
The next cycle, the first 1-6 non-overlapping valid instructions are accepted into a queue for further decoding. The NOP almost certainly takes up space in this queue.
At no point does RIP get incremented by one. There isn't even a single physical RIP register to increment, the CPU is "executing" dozens or even hundreds of RIPs in parallel.
dlcarrier | 3 hours ago
fluoridation | 3 hours ago
EvanAnderson | 5 hours ago
JoeAltmaier | 5 hours ago
mito88 | 8 hours ago
Score: 1 cycles Time: 0 nanoseconds
layer8 | 8 hours ago
hyperhello | 2 hours ago
Retr0id | 9 hours ago
jonathrg | 7 hours ago
twothreeone | 4 hours ago
Brian_K_White | 3 hours ago
michalsustr | 9 hours ago
inigyou | 9 hours ago
rrampage | 9 hours ago
pbsd | 7 hours ago
IshKebab | 8 hours ago
It would be much more interesting to know the results if you're only allowed to use main memory.
markus_zhang | 8 hours ago
monocasa | 7 hours ago
> Trapped/emulated/virtualized instructions may only time the trap, not the handler.
But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.
simonebrunozzi | 6 hours ago
[0]: https://en.wikipedia.org/wiki/Core_War
eek2121 | 6 hours ago
baddash | 6 hours ago
darksim905 | 2 hours ago
john_strinlai | 2 hours ago
kazinator | 2 hours ago
E.g. we can build a board around a MC68000 where we make it lock up forever in a bus cycle, waiting for a DTACK that doesn't arrive.
Some early microprocessors had clocked bus cycles without handshaking. They would put out an address on some address lines and signal some line together with a read/write indication, and then expect the transfer to be completed within some clock cycles. If nothing is attached to the address, they would read whatever values are on the bus, like maybe all 1's if it is an open drain system that requires the transmitting device to pull to ground to indicate zero.
I'd say that kind of thing belongs to a hall of shame; it requires software hacks to interface with anything that can't keep up with the prescribed bus cycle.
Joel_Mckay | 18 minutes ago
https://en.wikipedia.org/wiki/Metastability_(electronics)
Ultimately... failure modes must arise because naive gate design simulation models can't determine issues in a computationally feasible time frame.
The consequences of "fixing" CDC prone design flaws makes a processor many times slower (8 to 16 times slower on my dumb attempt), and develops weird alien design features very different from Von Neumann architectures.
This is why we can't have nice things. =3
56767865678 | 34 minutes ago