korrents

korrents · piece

Multimodality and Large Multimodal Models (LMMs)

Chip Huyen · 10 Oct 2023 · huyenchip.com

3 korrents from this piece

Chip Huyen did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. A model that can effectively learn from bitstrings or bytestrings will be very powerful, and it can learn from any data mode.
  2. Image is perhaps the most versatile format for model inputs, as it can be used to represent text, tabular data, audio, and to some extent, videos. There's also so much more visual data than text data.
  3. Text is a much more powerful mode for model outputs. A model that can generate images can only be used for image generation, whereas a model that can generate text can be used for many tasks: summarization, translation, reasoning, question answering, etc.