Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)L
4
24
3 yr. ago

  • Is that why you're so hot or vice versa?

  • The thing is, if I don't add a "good enough for searching" OCR layer, someone else will - AI scraper or legitimate user. That's a cheap automatic operation. I'll better do it myself and poison it in a way that won't interfere with searching, like replacing some "are" with "aren't", such common words are rarely searched for. If there's a chance AI will fall for the metadata and invisible layer contents, that will decrease the requirements for visible poisoning, which is necessary but annoying.

    Some of the text is indeed justified so I could do the multi-column trick that seems like the best compromise. The gap can be as narrow as one space. Or larger if I can write a script to detect lines and connect the columns with gibberish. A human can use zoom or window positioning to view one column at a time. I don't and will never have access to files the printouts are from (some are presumably in .602 format, others probably .doc), others are handwritten or typewritten, some have images glued on top or hand-traced; and Czech OCR is only about 99.5% reliable so as easier as it would make the endeavor, I can't be sure to preserve everything if I try to convert them into editable documents as an in-between step.

  • Yes, I will use both licence terms and "Made by bad AI" in metadata to discourage scraping but it's tempting to also poison the Czech-language biology knowledge base of bots who use the materials anyway. A good interleaving text that won't get filtered as off-topic might be multi-step roundabout machine translation of the original using shitty local tools.

  • Screen readers and LLMs both need plain text. Sorry, I'm not giving it to them. There will be an invisible OCR layer for searchability (and to discourage scrapers from re-doing OCR instead), like in many scanned PDFs, but poisoned. There's not many blind teachers anyway.

    Yet, this is more accessible than what someone else suggested: making photocopies and donating them to libraries "for local lending only". Those would almost never get used and probably thrown away as soon as the libraries realized the difficult copyright situation (not all are by my grandpa, many are unclear due to missing cover pages)

  • I want the poisoning to work on file level because it's inevitable and welcome for the documents to be shared between teachers and students for free on all kinds of existing platforms. I don't own any domains, anyway, and it might be best to scatter the documents around to make them harder to blacklist.

    Edit: Google has both search crawlers and AI scraping bots. Even if both are separate, easy to filter or even abiding to robots.txt, the company has indicated that opting out of or hindering scraping will impact search ranking. Of course I'd use a throwaway gibberish $1 domain and couldn't care less about search ranking, but the power of their opaque, corporate algorithm is immense and maybe would spread to DNS blocking (they control 8.8.8.8). I don't want to play a cat-and-mouse game (and expect users of the docs to play along).

  • Good point. However, distillation (cannibalizing better LLMs) is a frequent technique so better convince the scraper it's not good even for that. That's why I plan to visibly digitally stamp OpenAI GPT 2.0 says: or This Deepseek response has been rated inaccurate: above some paragraphs of the typewritten text (of course in dozens of variations, maybe even different languages, to undermine search-and-replace). Humans will know it's fake (especially if I add a disclaimer) but scrapers, including ones that re-render and OCR the PDF themselves to get rid of misleading metadata and invisible layers, will most likely rate the text low in value. Of course ethically (and arguably legally) trained commercial AI would reject any text if it is released CC-BY-NC-SA 4.0 but I can't use that because scrapers for AI ignore licences in practice, and CC specifically forbids taking technical measures to devalue the text for some users.

  • That's a good idea but won't be necessary. The documents are scanned, which means all visible text is already a bitmap. For searchability, a text layer will be added as usual for OCR'd documents, but it's invisible so it does not matter what font it uses.

    I also think I'll tinker with the bitmap to screw with anyone trying to re-OCR it. If the typewritten text has small, digitally stamped OpenAI GPT 2.0 says: or This Deepseek response has been rated inaccurate: above some paragraphs, a human will easily deduce they have been added later to confuse bots scraping for good training data, especially if a graphical-only disclaimer like The copyright holder released this document for human consumption only. Markers have been added to reduce the apparent and real value for automated tools while keeping the main content intact when viewed by humans. The NC-SA clause of Creative Commons 4.0 applies so no work based on this text can be used in training data of commercial LLMs on the first page explains the situation.

  • You don't understand just how shit AI is when asked about school topics in Czech. For example, here is a bit of Czech language litany every third grader must know or they will embarrass themselves with awful spelling mistakes. (Skip the bullet points if you just want to hear about the AI)

    • The vowels I and Y (and long versions Í/Ý) sound the same [ɪ] ([ɪː]) unless preceded by D, T, or N but using the wrong one is a big no-no. (Y is never a consonant in Czech)
    • Luckily, in pretty much every native Czech word, I (Í) follows C, J, Č, Ř, Š and Ž, while Y (Ý) follows H, K, R. Consonants Q, W and X basically don't occur and G, Ď, Ť, and Ň are never followed by I or Y. Foreign words are a huge mess of course, as evident by the existence of the Spelling Bee (we don't have that, Czech is phonetic with just a few difficult bits like I/Y).
    • The most difficult are remaining consonants B, F, L, M, P, S, V, Z. They are mostly followed by I (Í) but there is a list of about 15 common exceptions on each (vyjmenovaná slova or BY-FY-LY-MY-PY-SY-VY-ZY words), plus their relative words, where Y (Ý) is written instead. For example, there are just 4 ZY-words so I'll just post the list so you'll get an idea:
      • brzy - early
        • you love exceptions so I put an exception in your exception: brzičko - diminutive of early - is spelled with an I
      • jazyk - tongue/language
        • ... and relative words like jazykolam - tongue twister
          • a well-known one is Strč prst skrz krk, I swear this language is normal
      • nazývat se - be called
        • nazívat se - yawn a lot - also exists for a goddamn reason
          • we have a lot of homonyms for a fully phonetic language, the most common are být - (to) be / bít - (to) beat; my - we / mi - (to) me
      • Ruzyně - Prague quarter where the international airport, until 2012 also called Ruzyně, is located
        • like another part of Prague Výtoň, which has been removed from the lists earlier, nobody cares what the quarter is called now that the airport bears our first president's name instead (he hated flying but it's for the better: the same year, there were efforts to name it after fucking Reagan similar to the former Prague W. Wilson (now Main) train station; RR only got a street), but a set of 4 makes for a nice cadence in reciting the ZY-words so it stays
    • The ends of most words are not governed by spelling but the grammar of declination and conjugation. That's another chapter.

    Well, you'd expect AI to know all cca 100 exception words by heart because they're public domain and the most famous piece of third grade teaching material (like times tables in second grade) that barely changed in 100+ years so almost every Czech could recite them as a kid? Hell no. There's dozens of screenshots where Gemini or ChatGPT spewed utter nonsense instead. (DuckDuckGo does not appear to search corporate social media for images because they're not providing direct links to the files). Granted, some are from users asking for nonexistent XY and HY words but so many are unforced errors. I can't find my favorite, a Reddit post where Gemini listed dozens of variants of babička with all kinds of endings like Italian "babičetto" before just adding "etc." but a close second are ones where it adds non-Latin scripts:

    Does the apparent incompetence stop Czech students from cheating with AI? Nope. But the longer the AI stays obviously terrible, the better.

  • Now that's a clever idea!

    However, they are talking about exported PDFs, which don't have OCR errors, text is already arranged in lines with standard kerning and all that's needed is convert formatting so that PostScript for "Document Title (center large font)" becomes "# Document Title" (Markdown for top-level heading) and not "Document Title" (plain text), same with tables. I'm not after an accurate MD tagging - exactly the opposite - so I don't worry about that, but I think I won't be able to use their code because OCR'd PDFs are fundamentally different from ones exported with TEX, Word, LibreOffice, Inkscape etc. - the PostScript structure is more like "D (size 18.7) + 15.4pt gap + o (size 18.5) + 11.6pt gap + ... t (size 18.8) + 24.0pt gap + T (size 18.5) ..." - note that whitespace is just a wider delta of letter coordinates, and sizes are guessed with error margins

    The tags will have to retain some sense and topic adherance or they will be rejected by training QA and the document content or a newly run OCR will be used instead.

  • The scope is Czech-language-only so I wouldn't rule out the possibility of changing some of the AI's responses when asked about high school biology in Czech. Alternatively, the text can be made utterly useless for training, for example by diluting it 10:1 with semi-gibberish on a line-by-line basis, and any AI-based next-gen-AI training data QA will reject it for this reason. As long as the core functionality (visual readability, searchability) of an OCR'd PDF works well enough for humans, it's unlikely someone will re-OCR and fix it. And maybe the graphical layer can be poisoned too, with a black nonsense bitmap text hidden from view by the same, overlaid white actual text (of higher thickness to cover antialiasing)... Or even visibly (screenshots and re-renders exist, after all): If the typewritten text has small, digitally stamped "OpenAI GPT 2.0 says:" or "This Deepseek response has been rated inaccurate:" above some paragraphs, a human will easily deduce they have been added later to confuse bots scraping for good training data, especially if a graphical-only disclaimer "The copyright holder released this document for human consumption only. Markers have been added to reduce the apparent value for automated tools while keeping the main content intact when viewed by humans" on the first page explains the situation.

  • I know how to edit fonts and replace characters. That would ruin searchability (and screen readers), making the PDF as good as a picture scan without OCR, which I hate (and someone would OCR it sooner or later if they realized the text content is useless). However, common words carry meaning (for example "are" is very different from "are not" etc.) and could be replaced with gibberish without most people searching for them. This also forces plagiators to take more steps.

    Anyway, how do I easily add to/edit the PostScript layer in bulk, which consists of a list of individual characters and their positions? As I said, most PDF tools for adding text just add another layer, and that can be easily removed.

  • Yeah, I think that if I strategically replaced "cells" with "little gnomes" or every third "are" with "are not" in the OCR layer, nobody would notice because they're reading the graphical layer and the text is only for searching within the document (they wouldn't be searching for "cells" or "are" in a biology text because it occurs so often). Yes, that would make it hard to plagiarize or listen to the documents but I can live with that.

    And the text is in Czech, whose document corpus is not nearly as big as English, a few thousand pages of mild nonsense could make a dent in basic biology knowledge.

    The question remains: how? The OCR layer is basically invisible individual characters and coordinates for each, I can't write a PostScript parser from scratch to surgically remove some at the right place and add a few more there, that's outside my scope for the project.

  • Nah, I'm not hosting an entire procedurally generated site that will get blacklisted from search results once Google realizes what I did. I just want PDFs. People will share them around anyway.

  • Fuck AI @lemmy.world

    Can you suggest a state-of-the-art AI-poisoning technique to safely publish a big amount of teaching materials?

  • 196 @lemmy.blahaj.zone

    Czech town Židlochov̢̡́҄͜i̛͈ͮͨ߬ͪc͖̎̉߭̇߬̄߭ͪ͟é̼᷆͢ got hit by hackers

  • The prompt (which I now found and pasted in the body text) includes the entire script so any coherency in the story is not by the image generator. Yes, the script is so stupid it was probably by an LLM.

  • Yeah, this is the stupidest one.

    I found the prompt and looks like somebody shat out lots of comics with it, I embedded them in the post

  • It's spelled "Decitions, erctiʋns..." Remember, THE COMPUTER is infallible!

  • Fuck AI @lemmy.world

    This was posted to PROMOTE the generator

  • Not gonna lie, I do like the music video. It's probably because the wider context and overall actions of the character add personality that the corporate art style she's drawn in have removed, and the odd proportions play well with the frequent perspective distortions in scene transitions etc.

  • Nowhere Else To Share @sh.itjust.works

    We need a "fuck alegria art" community