Async Rust: Where does the scheduler live?

57 points by mond a day ago on lobsters | 31 comments

notgull | a day ago

The complaint about Tokio having the "hidden runtime argument" passed around is very valid. I think the attempts to "erase" asynchronous runtimes and attempting to make them invisible to the user hurts more than it helps for comprehension.

[OP] mond | a day ago

Yeah, that's where I arrived. I don't know if I'm right, but that's the one thing that just doesn't quite sit right with me.

This isn't unique to Tokio.

Any sane async implementation (i.e. one which can run without threads at least as long as you're not doing async file IO) is going to have some global/thread-local (potentially even scoped in some intriguing way) state in rust.

That's not to say that this design decision is great, but there are also plenty of downsides to the alternatives.

withoutboats | 8 hours ago

That's not true, you could pass the executor and reactor handle(s) down instead of having APIs like task::spawn or TcpStream::connect. Similarly you could pass an allocator to every API that allocates. It's a trade off; one unfortunate effect of choosing these APIs and not shipping an analogy to #[global_allocator] for the executor/reactor is the tokio monoculture.

Of course any scheduler is going to have state (which is the thrust of this post) but it doesn't have to be non-local state, you could pass it around by reference.

std::task::Context can't hold these references, which means that your executor can't pass them to a future it wants to poll, which means that the future has to be given this reference at construction time, which means that everything has to take a runtime reference and pass it along. This would be insufferable.

I mean what I said, having every single async function in your entire codebase gain a runtime reference parameter is not a sane option.

The current design of rust async prevents any other sane option. A world does exist where rust let you put that reference in the context, where you'd maybe expect it to be. Funnily enough, as I just learned, there's some effort to make such a world happen: https://github.com/rust-lang/rust/issues/123392

It has the downsides I alluded to, but it looks less terrible than I actually had imagined.

withoutboats | 4 hours ago

Not every future, just futures that mutate the scheduler state in a manner other than waking a task. This means things that spawn new tasks or register new event sources on the reactor, like the examples I gave. This isn’t my preference either but it’s not insane.

Indeed, it’s done that way in embedded Rust (bingo!) by passing a Spawner around. This is no hardship in my experience, perhaps because in embedded one typically has a pretty static set of tasks, most or all of which are set up at init time.

(Edit: And in my embedded code I typically make all the executors explicitly, as well.)

This is because in embedded, the executor doesn't have to call the wakers. Wakers are called directly from interrupt handlers. The reason it doesn't get out of hand in embedded is precisely because the runtime doesn't need to be passed to the future that it's spawning, because that future doesn't need to mutate it to set up wakers.

edit:

But the interrupt handlers are the global state.

Also, when it comes to sleeping, you again end up with global sleep handling state which talks to a timer queue which is state for an interrupt handler which dispatches wakers...

It's easier to separate in embedded, but if you were dealing with some kind of abstraction which had poll or something, you'd be back to global runtime state.

The hard part is not the spawning, and it really doesn't matter if that's global or in your face, but since stuff already has to be global to work correctly, those spawners may as well be too, because you can re-use the same tracking infrastructure to figure out which spawner you want to use since in userland rust you need to keep the waker state together with the spawner state.

Of course there is global state to make it work. The question was whether you access that state through implicit global “magic” or through explicit references.

Embassy does it both ways: explicit spawner.spawn() rather than task::spawn(), macros to “magically” spawn static tasks if you’d rather, and implicit Timer::after() using the “magic” time driver. I think smol has a similar split (global async-io reactor, explicit Executor) though not sure because I haven’t used it.

So yeah, it’s easier to justify global reactors (as opposed to spawners) since they’re usually reacting to something outside the execution model anyway, like an interrupt.

The way it works in embedded rust is that you literally access that state through magic memory addresses that you write to.

Again, stop focusing on spawning. I encourage you to actually write a basic rust runtime without threads which implements sleep and/or socket IO. You will learn a lot and understand the problem.

Since you're talking about embedded, you should also look at how embassy implements async drivers. Because I know them very intimately and can assure you that there is unsafe globL magic as opposed to regular global magic in embedded (for both sleeps and for async drivers).

How will those futures get that reference if their callers don't hold it?

Do you suggest it is made global and passed at the last moment? What benefit does that serve?

edit: To clarify, the hard part isn't the spawning part, it's having all those deeply nested futures know who to talk to in order to register their wakers.

withoutboats | 2 hours ago

No but most futures don’t spawn other tasks or create IO objects. Zig does actually pass a handle around to do these things (not to mention an allocator), it’s not insane.

Most futures only need to interact with the scheduler by scheduling more work: that is what wake does.

Futures do not generally “register their wakers.” Futures are passed a waker when they are polled.

I don't want this to sound rude but you should actually impement a basic rust runtime which implements sleep and/or poll/epoll/select based socket IO so you can understand the problem space (you are not allowed to use any kind of thread).

You cannot do anything useful from a future without access to the runtime from inside a future, and you cannot access the runtime from inside the future without a reference to it, the context only contains the waker, and nothing else.

Zig does actually pass a handle around to do these things (not to mention an allocator), it’s not insane.

That's cool, I have no idea how rust async works but presumably, like async in many other languages, it gives futures some channel back to the runtime which doesn't require global state or passing references through every async function.

Like I said, it's possible, but it has its own drawbacks for rust, specifically in regards to the type system.

See the tracking issue I linked for details.

Futures do not generally “register their wakers.” Futures are passed a waker when they are polled.

And why do they get passed a waker?

It's most assuredly not to call it/not call it directly. That's useful for only a tiny subset of special case futures.

They are passed it so they can give it to someone who will call it when the thing they're blocked on to make progress. And in all useful non-embedded async contexts the thing they pass it to is the runtime. So unless your futures do not actually eventually do anything meaningful, they must be able to communicate their needs to the runtime.

Any sane async implementation (i.e. one which can run without threads at least as long as you're not doing async file IO) is going to have some global/thread-local (potentially even scoped in some intriguing way) state in rust.

Why is that?

At the end of the day, any await on a future will eventually end up at something which manually implements Future. Future has one method poll which takes a Context this context only contains a waker at the moment, and it's not generic in any way.

So you have two options, have every single async function in your entire program take a runtime reference parameter which is passed whenever constructing a future, or use globals.

You avoid function coloring. It’s a function parameter.

I have a problem with this wording, because function coloring is isomorphic to function parameters. It is fundamentally about how if you need thing X as a parameter at the bottom of the callstack, it will spread upwards. It doesn't really matter whether the parameter is implicit or not. And it also goes for return types: if you introduce a Result at the bottom of your previously-infallible stack, it will start to spread upwards, unless you bite the bullet and unwrap it. Which will then unwind the stack so philisiphically it still does spread upwards, except implicitly and invisibly.

carlana | a day ago

I think the bingo card angle is very funny.

And this particular bingo card is so damn accurate.

epilys | 14 hours ago

akavel | 6 hours ago

Wow, that's a great post, I learned a lot about async Rust from it that I had no idea about! Especially from patiently reading all the collapsed sections 😁

Regarding the passing of IO/async/... in Zig - I don't know much about it, but just from what you wrote, it also sounds interesting to me from a point of view of an explicit "capabilities passing/tracking" system! Or, in other words, "side-effects" stuff, like in Haskell I think maybe? (Also don't know that one well.)

easrng | a day ago

is it just me or does this page crash firefox

proctrap | a day ago

seems to be just you

[OP] mond | a day ago

Might just be you. I mainline Firefox, and don't have this issue.

Waterfox is also not causing any issues.

valpackett | 3 hours ago

The crate thing is how you get Tokio defaultism, and a lack of runtime agnosticism

If only there was some kind of way for all libraries to be parameterized over the runtime…

I don’t have it in me to do research on that today, but figuring out which types of programs would benefit from [kernel-managed M:N thread type things] and what the limitations are is an interesting question. In practice, the kernel will never know as much about your code as a language-specific runtime.

With the "kernel vs userspace for $thing" question, it's faster to "stay where you already are". Context switches are expensive. And it seems really hard/awkward to schedule from the kernel without yield points being extra context switches. With old-skool poll-based I/O, you could maybe try to attach the next task choice onto the poll syscall itself but that would only improve the situation for the tasks that are directly waken up by external events (like socket readiness). Bouncing from one task to another would still be "extra" poll calls then.

Now that Linux has realized that completion-based I/O scales a lot better, we're firmly in the "fewer syscalls actually" territory…

aw1621107 | 2 hours ago

If only there was some kind of way for all libraries to be parameterized over the runtime

Fun fact: Graydon Hoare wanted Rust to have ML-like first-class modules!:

Traits. I generally don't like traits / typeclasses. To my eyes they're too type-directed, too global, too much like programming and debugging a distributed set of invisible inference rules, and produce too much coupling between libraries in terms of what is or isn't allowed to maintain coherence. I wanted (and got part way into building) a first class module system in the ML tradition. Many team members objected because these systems are more verbose, often painfully so, and I lost the argument. But I still don't like the result, and I would probably have backed the experiment out "if I'd been BDFL".

Makes me curious how such a Rust would have turned out, especially if some of the more modern OCaml module features mentioned in your link were implemented.

alper | 12 hours ago

trying my best to complain about async Rust, only to end up going “Oh. Yeah. That’s reasonable. That makes sense. I can see why they did it that way.” every step of the way.

I never understood the complaining about async rust. Is it like kubernetes/react where people just kneejerk into it?

nytpu | 7 hours ago

Because they demonstrably—and self-admittedly—hurriedly standardized a clearly (with hindsight) flawed minimum viable product; that, while usable, could demonstrably be improved in some ways (even if many of the complaints people voice about it were valid tradeoffs), and then proceeded to do almost nothing to address any of the issues until eight years later in 2025?

e.g. Pin was them slapping together "whatever they could think of" because they couldn't figure out how to make a Move trait work at the time, and now they're stuck with Pin forever for backwards compatibility. And on top they then proceeded to leave pinning as-is with zero changes to make it more ergonomic: it's a pure library type with no syntactical sugar, there's no automatic reborrows, there's no compiler help to enable the use of pin projections (and the hacks to enable pin projection have massive caveats), it's difficult to reason about, etc. etc.

ekuber | 6 hours ago

You're making it sound like the people in the project were twiddling their thumbs and not, well, doing things, including in that space.

https://github.com/rust-lang/rust/graphs/commit-activity

https://github.com/rust-lang/rust/pulse?period=monthly

For comparison, as a first approximation, the metrics of the entirety of my work in the last decade fits in a month of work of just the compiler.

I know that the work for Return Position Impl Trait, which was needed for Async Functions In Traits, required a years of focused refactoring to land. Similar things with const evaluation. There's so much work happening behind the scenes on so many diverse domains that I don't think there's a single person that knows all of it.

I often am frustrated with the speed of the project. The same carefulness that makes it easy to update the toolchain every 6 weeks means that we collectively have an unwillingness to land 80% solutions if they have the chance of causing churn. The async fn in traits from above is a good example: we could have landed a Box based solution half a decade ago, which no one would be happy with but that would have works and could have been migrated to the current approach. But, amongst other considerations, I suspect that plugging that feature hole would have removed the wind from the sails of TAIT and RPITIT, meaning that today we're in a better situation than we would have been. Specially because the stop gap was a single annotation on a trait, a language wart, long in the tooth, that now can be removed. We could have landed generator expressions and items 5 years ago, but those conversations stalled on answering questions about async generators, consuming them, and trait AsyncIterator/Stream. Different people are working on that now, making progress.

You have to understand that Rust is not Swift: we don't have teams of employed people working on deliverables where if another team is blocking you you can raise a stink to get things moving, nor is the "product" one that can have breaking changes every few years. The user base is incredibly diverse, from embedded to high performance distributes systems, and everything in between, and the project has a distaste for any feature that would make things better for one by making them worse for another.

Also, Rust is a showupocracy. The things that get done are not necessarily the ones that are clamored for in a forum, but rather the ones that people show up to do. Sometimes those that show up do so for their employer, others it is a personal pet peeve. The best way of making sure something is done is to do it. But you also need to be aware that the answer for "why don't they just do X?" is a well reasoned, infinitely argued over, 50 page treatise on type theory, ergonomics and platform support. Many of the things that don't land quickly, is because there are open questions. Answering those satisfactorily is a way of speeding things up.

mwcampbell | 6 hours ago

Thank you for that explanation. Thanks also to @withoutboats (who also commented on this thread) for doing so much of the original work on async support, and for continuing to participate in some of these discussions, despite all of the complaints and criticisms. FWIW, Rust continues to be my go-to language for new projects, warts and all.

nytpu | 5 hours ago

Yeah, that's all true, I was really unreasonably hostile in my reply, sorry…