Ars Astronomica – English translations of rare Hebrew and Latin astronomy texts

102 points by sweisman 15 hours ago on hackernews | 44 comments

[OP] sweisman | 15 hours ago

I built Ars Astronomica to make historically important works on astronomy, calendrics, and related sciences accessible in modern English. The project combines OCR, AI-assisted translation, editorial review, and publication into a workflow that can handle book-length works while preserving diagrams, mathematical notation, and references.

The site currently includes English translations of rare Hebrew and Latin works, including texts that have never previously appeared in English and others that survive only in manuscript form. All editions are released under a Creative Commons license.

I'm interested in feedback from historians of science, classicists, medievalists, digital-humanities researchers, and anyone interested in AI-assisted scholarly publishing.

pmontra | 9 hours ago

I don't belong to any of those categories of people but I'm glad that you took the effort to build this site. I didn't know about those texts and I will at least skim through a few of them now. Thanks.

madhu_ghalame | 13 hours ago

This is a valuable initiative. It would be great to include more historical context alongside each translated work.

domschl | 13 hours ago

Why was the comment of the author of the translation-site that explains the methodology flagged / dead?

The mere fact that the translation was AI-assisted shouldn't be a reason to flag a comment. AI translations, especially supervised translations, can actually achieve a reasonable level of quality.

Disagreement on AI philosophy shouldn't be a reason to flag a contribution.

[OP] sweisman | 13 hours ago

A professional book editor and Hebrew translator said the translation of the one work he checked out was "impressive".

Others have responded equally positively to the Latin translations.

AI as a panacea is overwrought, but there are genuine positive uses of it. Would critics rather have NO translations of these pivotal works?

NetOpWibby | 13 hours ago

Yes, because if a human didn’t do it, how will we know of the translation is free of mistakes? /s

[OP] sweisman | 13 hours ago

Professional translators don't throw around the word "impressive" like that.

You don't, yet I still stand by my work.

I translated some books up to four times as I refined the process. The default behavior is repeatedly to not attempt to translate text it can't resolve for whatever reason. AIs do have a tendency to hallucinate, and I expended a lot of effort on minimizing this problem to the point it can be considered mitigated.

ButlerianJihad | 8 hours ago

Professional publishers collect blurbs which are attached to real names of real people, and bona fide reviewers write in full sentences with context, without

[OP] sweisman | 8 hours ago

It's good enough for me. I'm not publicizing his name without permission. He is quite well-known and until I moved recently, was a neighbor of mine.
> Would critics rather have NO translations of these pivotal works?

I sent a Classicist the link because it's adjacent to his area of interest, and by chance he addressed this question. I don't think he'll mind me quoting him:

"There's a lot of this shite popping up at the moment. [...] There seems to be some idea that publishing a shit translation of an untranslated work is better than no translation, but that's obviously wrong."

Speaking for myself: Some of the anomaly rules are obviously patches to deal with a specific situation, and might be better off in code. But relying on the chatbot to flag wider anomalies is odd. How does it know what it doesn't know?

At a minimum, I'd use a spread of models to gain binocular vision, and I wouldn't publish until a human was prepared to sign their reputation to it.

In other words, I think what you've got there is a first draft of a translation, not a translation. Given that, if you are going to publish, I think the NoDerivatives restriction is a mistake. But that's a minor issue.

[OP] sweisman | 6 hours ago

So, he's commenting on something he didn't read and compares it to shit? Good to know. I'll ignore him. I understand his point. Still, obviously my rhetorical question had in mind that what I do is good enough quality to make the question worthwhile. Better than a "first draft". As mentioned in another comment, some works were translated up to four times as I refined the process.
The thing with LLM-generated code is you can test it on a real CPU - that's a tight feedback loop with a cast-iron verification mechanism.

Translations of human language have no analogous mechanism. So yes, the translation might be perfect, but until a human puts their reputation on the line and says "I certify this translation is accurate", it's still shit. It's shit because of the way it was created, not its absolute accuracy (or otherwise). He doesn't need to read it.

I'm sorry, I'm not trying to upset you, and I know I won't change your mind. We just have different philosophical positions on this, I think.

[OP] sweisman | 2 hours ago

I know what I know from experience. You know what you know from nothing at all.

Numerous people have commented to me about Hebrew and Latin works, who know what they are talking about, and none has criticized the fidelity of the translation to the source. The criticisms proffered are minor, while the praise extensive.

A couple years ago, I shared your opinion. Now I don't. You are the one who won't change your mind. Like you said, it's "philosophical" for you, while for me, it isn't. If you insist on calling it shit regardless, you are retarded.

> You know what you know from nothing at all.

I've done it myself. Transcription of 19th century newspaper articles to markdown, mostly. Some earlier wills (which were an absolute pig - secretary hand). Oh, and categorisation of postcards. That's why I was poking around your pipeline - to see if I could learn anything. You're right, tabular data is hard. Also columns, and proper nouns.

Feeding the LLM a context-aware cheat sheet helped with the nouns. BTW, what I said about using multiple models for parallax was good advice.

Do you know you're very spiky?

names_are_hard | 6 hours ago

I'm not a book editor or translator, just a guy with a reasonable familiarity with rabbinic literature in its original language. I spent some time reading the translation of Tzemach David, and I can say the following: it's legible to me but the AI is clearly lacking a style guide. It's inconsistent about what it chooses to translate and what it chooses to transliterate, and when it transliterates, it's inconsistent in how it does so. Eg "Rosh mesivta" vs "metivta" vs "exilarch". The AI is drawing on a hodge-podge of training materials with diverse audiences and no clear audience of its own.

[OP] sweisman | 2 hours ago

Can you email me with more info? It is likely the distinctions you noted are from the source. There is a glossary maintained that keeps specific words or phrases consistent. But slightly different phrases can resolve in inconsistent ways.

The problem is not the AI, but rather a weakness in the design. In other words, a bug. Or a regression and we call them today. It is fixable.

kensai | 12 hours ago

"Disagreement on AI philosophy shouldn't be a reason to flag a contribution."

Amen!

bawolff | 6 hours ago

> AI translations, especially supervised translations, can actually achieve a reasonable level of quality.

That may be true, but also not really what i'm worried about. With AI on an ancient book written from a world view very different from the contemporary one, i'd worry more about misleading translations.

[OP] sweisman | 2 hours ago

Misleading in what way?

bflesch | 13 hours ago

Nice project, but I think the PDF format could be improved. I'd prefer having landscape layout and original text and translated text side-by-side or keep the portrait layout and do a line-by-line translation so that it is easier to review and spot translation errors.

[OP] sweisman | 13 hours ago

Thank you.

Your suggestion makes sense, and you are not the first to offer it.

Maybe someday, although probably not CC licensed.

lordgrenville | 9 hours ago

While it's really cool that this is possible now, this project seems to be a thin wrapper around running Claude. Presumably, someone who wanted to read one of these books could just run it through Claude themselves with comparable results.

Sharing some artifact of your own added value, or things you learned in the process, would be more interesting to me than the output.

us-merul | 9 hours ago

I think there's more to it than that. It's a combination of extracting the text from the PDFs, resolving persistently tricky issues (especially numbers), and flagging issues for human review at scale. You can view lessons from the author on the GitHub repo: https://github.com/sweisman/translation-pipeline

lordgrenville | 9 hours ago

Great, then rather post that! Although that repo looks like 100% AI artifact, not human intervention. (What does "flagging issues for human review at scale" mean in this context?)

us-merul | 8 hours ago

I was referring to this bullet under Lessons Learned: "Manuscript translation is feasible but rough. A 14th-century Sephardic manuscript hand is readable at ~70–80% accuracy. The output is a useful rough draft, not a polished translation. Plan for a longer discrepancy report and more editorial review." Some of the samples I looked at had question marks or original Hebrew text throughout, which I took to mean marked uncertainties.

[OP] sweisman | 8 hours ago

Manuscripts are doable but tougher than printed works. Hebrew likewise is tougher than Latin.

[OP] sweisman | 8 hours ago

This is not what I use to translate. It was the first version to create usable output, and intentionally implemented in a code-free manner (today I have an extensive body of code in addition to the Claude Code harness instructions). Read my origin story, linked to in another comment, for more background.

lordgrenville | 8 hours ago

Reiterating what I wrote above, sharing these would have been more interesting imo than sharing the pipeline output.

[OP] sweisman | 8 hours ago

I can assure you this is not a thin wrapper. Even the public repo of the first workable version of it is not a thin wrapper.

You are welcome to try. It is not rocket science, but what I developed would take much effort to reproduce.

I wrote an origin story about how I got started: https://aicentral.substack.com/p/pipeline-to-the-past

It describes some of the added value and lessons learned.

mmooss | 5 hours ago

Also, people can look at the Github repo of your prior version and see it's clearly not just 'running it through Claude'.

https://github.com/sweisman/translation-pipeline

What I haven't found in either the OP or Github repo is what your background is - software development? astrophysics? paleography? two or three of those? [edit: history of science?] - and what human review is involved. (My apologies if I overlooked anything; I don't have time to read it all!)

[OP] sweisman | 2 hours ago

Software development. I did work in an astrophysics lab decades ago, as one of my public repos documents.

dwohnitmok | 8 hours ago

"All translations in this collection are © Scott Weisman. All rights reserved, except as granted by the license below."

Does copyright actually belong to Scott Weisman if all the words are outputs of LLMs? Is there any relevant case law here? I'm very curious about whether LLM outputs are copyrightable by the LLM user (I'm guessing potentially it varies throughout the world).

[OP] sweisman | 8 hours ago

Yes. The rules vary worldwide. Do your own research. I did, before asserting copyright.

therealpygon | 7 hours ago

Solely the direct output of an LLM? Not in the US, no, according to the copyright office. Can a human edited and corrected text in which a human adds authorship (corrections, style, diction, editing, composition etc) to LLM output? Yes. Though simply making singular mechanical edits/corrections wouldn’t meet that bar; this seems like more than that.

An original human-authored translation is a derivative work that can hold copyright, but it is the human-authored parts that give it protection.

No amount of human authored pipelines that are automated (no human input) would give it this status as that is not human authorship (the pipeline itself can be however). The prompts for the LLM can also be copyright (again…human authorship), however they would be difficult to enforce since the direct output can’t be copyrighted.

tulio_ribeiro | 6 hours ago

Hey, man. Great work and thanks for sharing this as CC.

deluvenec | 5 hours ago

As a jew: it would be better to put up a site which does not risk to create the false impression of somehow comparable contribution of early european astronomers and medieval (and not so medieval) rabbis.

mmooss | 5 hours ago

Could you tell us more?: How are those groups not comparable in this context, and what do you see as the downside? And what is dishonest about the presentation?

deluvenec | 4 hours ago

I said that the contributions from early european astromomers and rabbis are not comparable, and that this site risks to create the impression that they are (I edited out the dishonesty part, as I can't really know what your intention was). If Copernicus and Kepler would be on this page, perhaps this would be clearer. But I really think that one Tyho Brahe suffices. His measurements, that is.

mmooss | 3 hours ago

I know what the prior comment said but I (and I imagine others) don't understand it. There is a lot of context or prior reasoning left out:

In what respect are they being compared, from your perspective? How do you think they are not comparable? What is the harm of comparing them? How would Copernicus and Kepler help and what is the drawback of multiple Brahe entries?

> I can't really know what your intention was

My intention? I have nothing to do with the OP. Many assumptions jumping up here.

[OP] sweisman | 2 hours ago

Judaism says revealed knowledge, as transmitted to us via our mesorah, is superior to empirical knowledge, such as that obtained via the scientific method.

That being said, Ralbag, as we refer to him, or Gersonides, as you do, was both a highly regarded (and controversial) rabbi and esteemed astronomer who pioneered empirical observation and measurement.

[OP] sweisman | 2 hours ago

Copernicus and Kepler have both been translated to English. I am focusing on works that haven't been.

Gans met and wrote about his encounters with Tycho. It was these documented encounters that drove my interest in translating the work.

In chapter 25 of the same book I mentioned, on the famous Gemara where Chaza"l concede to the chachmei goyim in Pesachim, Tycho says the Jews were right and the goyim wrong, and the Jews were wrong to concede.

Gans also met Kepler.

[OP] sweisman | 2 hours ago

David Gans, in Nechmad v'Na'im, the first book I translated, states as a matter of fact that any wisdom obtained by the nations regarding astronomy came from the Jews. Ralbag's treatise was so significant the Vatican commissioned a Latin translation.

In any case, my email is on the site and in every book. Feel free to contact me privately to continue the discussion.

mondainx | 5 hours ago

Don't get me wrong, this is awesome, but it could be better by adding Greek and Arabic, and possibly Sumerian documents?

[OP] sweisman | 2 hours ago

Thank you. I have considered Arabic, as there a couple works I am interested in.

My queue of Hebrew and Latin works is already lengthy, and this is my focus for now.