GNU social JP
  • FAQ
  • Login
GNU social JPは日本のGNU socialサーバーです。
Usage/ToS/admin/test/Pleroma FE
  • Public

    • Public
    • Network
    • Groups
    • Featured
    • Popular
    • People

Conversation

Notices

  1. Embed this notice
    Ulises 🌼 ⁂ /I\ (rataunderground@neopaquita.es)'s status on Monday, 13-Apr-2026 16:56:16 JST Ulises  🌼 ⁂ /I\ Ulises 🌼 ⁂ /I\

    "Hace falta un LLM open source".

    Vale, voy a atacar a fantasmas y hombres de paja porque yo sé que no es una postura con suficientes defensores como para que merezca ser tomada en consideración, pero el asunto es que no es posible un LLM open source. Los modelos no tienen código, no se escriben, y nadie sabe lo que hay dentro de cada uno o cómo funcionan internamente. Lo que se sabe es lo que se deduce por observación de lo que entra y lo que sale.

    Por ejemplo, los LLM entrenados para traducir de idioma A a idioma B y de idioma B a idioma C son capaces, sin que hayan sido entrenados para ello, de traducir de idioma A a idioma C.
    Esto implica que el modelo en su interior ha creado un cuarto idioma (D) al cual traduce cualquier input y después de ese idioma traduce al idioma de salida, pero no se puede saber cuál es este idioma D, totalmente encerrado en la caja negra.
    (Source: Enabling zero shot lenguaje translation https://arxiv.org/abs/1611.04558 )

    Puedes entrenar un LLM sólo con datos open source (¿qué licencia? No puedes mezclar MIT con GNU, por ejemplo) , puedes hacer open source el software para interactuar con el modelo, pero NO es posible hacer open source el modelo a menos que la definición de open source cambie, pero eso es mover la portería.

    In conversation about 6 months ago from neopaquita.es permalink

    Attachments

    1. Domain not in remote thumbnail source whitelist: arxiv.org
      Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation
      We propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no change in the model architecture from our base system but instead introduces an artificial token at the beginning of the input sentence to specify the required target language. The rest of the model, which includes encoder, decoder and attention, remains unchanged and is shared across all languages. Using a shared wordpiece vocabulary, our approach enables Multilingual NMT using a single model without any increase in parameters, which is significantly simpler than previous proposals for Multilingual NMT. Our method often improves the translation quality of all involved language pairs, even while keeping the total number of model parameters constant. On the WMT'14 benchmarks, a single multilingual model achieves comparable performance for English$\rightarrow$French and surpasses state-of-the-art results for English$\rightarrow$German. Similarly, a single multilingual model surpasses state-of-the-art results for French$\rightarrow$English and German$\rightarrow$English on WMT'14 and WMT'15 benchmarks respectively. On production corpora, multilingual models of up to twelve language pairs allow for better translation of many individual pairs. In addition to improving the translation quality of language pairs that the model was trained with, our models can also learn to perform implicit bridging between language pairs never seen explicitly during training, showing that transfer learning and zero-shot translation is possible for neural translation. Finally, we show analyses that hints at a universal interlingua representation in our models and show some interesting examples when mixing languages.

    Feeds

    • Activity Streams
    • RSS 2.0
    • Atom
    • Help
    • About
    • FAQ
    • TOS
    • Privacy
    • Source
    • Version
    • Contact

    GNU social JP is a social network, courtesy of GNU social JP管理人. It runs on GNU social, version 2.0.2-dev, available under the GNU Affero General Public License.

    Creative Commons Attribution 3.0 All GNU social JP content and data are available under the Creative Commons Attribution 3.0 license.