I was never quite sure why audiophiles went on and on about jitter. So I eagerly opened this thread.
My first thoughts.... If you have jitter to measure at x10^-6 seconds... you need a new clock. Actually, you probably need to stop what you are doing. Sell it all and take up a new hobby. Something is broken or you are doing it wrong.
I was bored so this happened.
48k 'clock' is meaningless.
Taking a bare, basic, $2 ADC and 16bit 48k stereo. You will hear all the talk about that 48k clock. In the hardware... wait for it... nothing runs at 48k. The 48k is entirely derived by other means. The one "signal" that 'runs' at 48k is the LR 'clock'. Honestly, it's not really a clock at all. It's the SPI SS line. Used in SPI - Serial Peripheral Interface to signal that a "slave" device is selected and the transmission relates to it, or that it is it's turn to send (bus mastering). Phillips just borrowed that signal line for the Left, Right signal. This is common, people borrow and modify SPI all the time, it's convenient. Later (Dolby) it gets used for "Channel strobe" to mark the first channel out of N. It's not used for reference by anything really and is mostly just used to make sure your left is your left and your right is your right. I mean "older" more blunt hardware "might" trigger the shift register to output the current sample. But really. It can count to 16 bits on it's own without a 48k signal telling it to, when in sync it will stay in sync and for the most part the LR clock is used a detect errors.
There is the serial bit clock (BCK or SCK). This is the real signal clock, or the primary clock if you will. At 16bit (non extended), this is 16bits * 2 * 48k (+ any guard space but we will leave that) = 1.536MHz.
Immediately you should see that a 1x10^-6 "jitter" in any of your digital audio gear is a very, very big deal. If you have that kind of jitter, it's not that your clock has jitter.... jitter would be like having miss aligned teeth.... yours is missing a few teeth entirely. Segway to another analogue, if your audio gear was actual gears, it'll be twitching and "jittering".
1.536Mhz isn't enough though.
However. While that 1.536Mhz clock can be generated and be a percent or two out and nobody will really notice, when you have multiple audio streams that MUST stay in sync, such as when recording or playing back multiple concurrent streams.... it's not accurate enough. 2 separately derived 1.5Mhz clocks can offset, skew and drift enough to be noticeable.
So the concept of the Master Clock. A clock with a far higher rate than needed, which can be distributed to "near by" equipment which shared synchronisation and... they derive all clocks and sample/trigger other clocks based on that master clock. Sometimes devices just use an internal clock derived off their even higher internal sysclock, such as in anything RPI and up and some MCUs/DSPs. Other times they will provide a master clock or require a master clock. 24.576Mhz is popular as it divides nicely into the 48/96/192 family of sample rates and their bit clocks. This is literally the clock which 'runs' the hardware. In digital NOTHING happens without a clock. It won't matter how many bits or samples go past that hardware if there is no master clock heart beat it misses it all. (suggested watching, Ben Eater YouTube).
Jitter at 24.576Mhz?
Jitter at 24.576Mhz is measured in nano seconds at worst. It IS there. The datasheet for the xtals et.al. have the figures printed with best and worst cases and usually a whole set of graphs showing the response under different VCC, temperature, load, capacitance etc. 50ppm or 100ppm might be quoted. I would caution that unless you read that whole datasheet and understand exactly what that 50ppm figure is for, you might find your accuracy nowhere near that in the wild. My breadboard clock shows a standard deviation of upto 50kHz! While that sounds like a lot, on such a high clock, it still only something like a 10ns jitter.
Clock domains, synchronisation and spacetime observer problem.
Jitter would not concern me there. Where it might concern me is keeping different clocks in sync, when required. When I asked the grey beards a few (I thought) simple questions on reducing the complexity in combining/crossing clock domains, they used words like "irreducible". Partly it comes down to distance and the speed of light in copper or fibre and even quantum mechanics as to just how irreducible the complexity is. 2 clocks, one either side of a room, 3 meters apart and 100 light nano-seconds apart. That's longer than the duty cycle of some master clocks. If your time reference is that clock, and each observer of that clock sees pulses at different times. Weirdness ensues. What about when the clock is more than 1 clock period in light further away from receiver than the bit clock source, whose clock is right? It gets worse when you start to have more than one pulse on the same copper or fibre. At this point we leave the 1970s and 1980s hardware behind. This is why people don't "ship" clock.
Where might you have issues. If it's not the digital clocks and clock rates themselves, where can the issue lie?
Noise. Parasitic capacitance, inductance, resistance. Environmental EMI. Special relativity, quantum mechanics. When you redraw the circuit diagram for your digital audio setup, include all of these parasitic elements when using LTSpice for example. Include all sources of noise. Add the speed of light delays on lines. You will find, as above, clocks don't ship well. Those lovely square waves don't just look like sinewave garage on your audio scope, they actually are now sine wave garbage. Sinewave garbage with reflections, resonances etc. The kind of sinewave garbage that makes the trigger on your scope twitch like it's had too much coffee this morning after a night out. Just like your DAC will. It becomes a game of "find the edge" and while most modern stuff will find that edge pretty accurately, there can be deviations on exactly how each device interrupts it. Not all are TTL, CMOS or even have published standards outside of the datasheet. Some us hardware edge triggers, some use super sampling shift registers and flip flop latches.
PCB Traces versus Einstein/Minkowski.
The general rule for sending Mhz across a PCB trace is to keep the distance not greater than 1/10th the light distance of the clock period. Light only travels 0.3m (1ft) in a nanosecond. I believe the 1/10th is to allow for any ringing to bounce back and forward a few times. So a 1Ghz clock shouldn't be expected to travel further than 30mm on a PCB trace. A 1Mhz clock you can get away with 30 meters! A 30Mhz clock would be about 1 meter. Of course there is a whole field of engineering butchered to death in that paragraph of over simplification .. and no I dont understand all the details of how differential pairing, shielding, compensation, regeneration, reconstruction etc makes a difference. I know the multi-ghz protocols don't use clocks, they modulate a bunch of waves into each other and send the data in the phase distortions instead. (again butchering differential and line encoding).
Note. It matters little if you use copper, fibre or open flat space vacuum. You can't go faster than light. People don't ship clock. Einstein screwed that up.
It's usually better to clock up. (clock!).
Where it matters, usually a sync pulse or strobe is used at a far, far lower rate than the actual bit clock. It's up to the devices to divide up that clock. It's usually more accurate to trigger a dividing PLL off a 1Mhz clock and generate 1.356Mhz from it (while that sounds like a particularly cruel example). This is how a CPU get a 5.8Ghz clock. By dividing a clock which usually starts out at only 200Mhz or less. As it gets closer and closer (in light pico seconds) to where the clock will be used it gets divided more. The same is occurring with the memory bus, the PCI(e) bus etc. etc. As an analogy .. a violinist does not need to only play when the drummer plays. The violinist can play 32nds or 64ths if they are good enough. 16th triplets maybe. I mean if that didn't work music would be a bit sh..
SPDIF, think ... green carpets and brown sofas, corduroy jackets and moustaches.
On SPDIF/ToSLink... my understanding is they use the same analogue modulation techniques to multiplex multiple digital streams together over the same signal line, that line isn't a just a digital pulse waveform. It has been a while since I delved into the details, but my understanding was, that modulation has holes in it, holes such that the incoming data cannot be encoded exactly and has to be modified. Such conditions exist that will not and cannot be detected at the other end and corrected resulting in audible distortion. It was considered "Good enough for the job" back then, when stuff ran at 16bit 48k. I imagine while it works fine as long as you can get a fast enough optical emitter/receiver and/or SPDIF chip, I also imagine those "don't ship clock" issues will still raise their head in the high Mhz there. It's 1970s tech for hell sake. The optocouplers/laser diodes got better, but still. I compand the two because I believe they are just the same line encoding sent electrically or optically. Just a different PHY layer. I might be wrong.
My advice?
Keep your synchronous digital audio connections short, high quality, well shielded and appropriate. For longer runs transfer to another standard, piggy back on a much higher bandwidth asynchronous serial protocol like HDMI (or USB heaven forbid) and borrow it's circuitry and encoding magics. Forget sync over more than the length of a room. The lower the same rate and bit depth the easier it will be to remain spacetime sane and avoid the vast majority of the issues. Ignore a dozen nano seconds, forgive rare microsecond glitch, fix it/replace it if it's got microsecond glitches all over the place.... or just stop trying to sync it and run a buffer.