One thing that I wonder (and maybe yall have answers to) is how much of the technical stuff in compute governance punt-able for when we get the political will/regime to have people seriously consider implementing them. What parts aren’t? What subset requires more time such that it makes it harder to punt?
I think if you're a technical person, it's worth helping on the issue now. The space over which politics can negotiate is strongly constrained by what is/n't technologically feasible. If you're a policy-but-not-tech person, then learning what is on/off the table politically and sharing with the techies is also valuable.
I was playing around with Inkling NVFP4 earlier today (mostly on some no-CoT time horizon evals). Inkling includes an MTP head for speculative decoding, so I was planning to take a stab at extending Token-DiFR and Activation-DiFR from Luke's paper to it. I am using Modal's sandboxed runtime, which works well for me, and doesn't currently expose DCGM counters, but it's possible that may change in the future.
incredible work- thank you both for writing this up
thank you for your terrific feedback!
yall COOKED
This was helpful.
One thing that I wonder (and maybe yall have answers to) is how much of the technical stuff in compute governance punt-able for when we get the political will/regime to have people seriously consider implementing them. What parts aren’t? What subset requires more time such that it makes it harder to punt?
a great question -- can imagine a version of this where we look at the calendar-time needed to implement the things
I think if you're a technical person, it's worth helping on the issue now. The space over which politics can negotiate is strongly constrained by what is/n't technologically feasible. If you're a policy-but-not-tech person, then learning what is on/off the table politically and sharing with the techies is also valuable.
Great post, saved.
I was playing around with Inkling NVFP4 earlier today (mostly on some no-CoT time horizon evals). Inkling includes an MTP head for speculative decoding, so I was planning to take a stab at extending Token-DiFR and Activation-DiFR from Luke's paper to it. I am using Modal's sandboxed runtime, which works well for me, and doesn't currently expose DCGM counters, but it's possible that may change in the future.
Nice work! My own small contribution to the discussion is here:
https://canaryinstitute.substack.com/p/if-you-cant-trust-then-verify