Mark2TeX is a LuaLaTeX package for Markdown-like text fragments inside
LaTeX documents. It reads .md files during LaTeX compilation, translates
the currently supported dialect to LaTeX, and inputs the generated .tex
file back into the document.
The focus is on a LaTeX-friendly writing style for scientific documents, not on full Markdown or CommonMark compatibility. If compatibility with common Markdown tools happens as a side effect, that is welcome, but it is not the primary project goal.
Mark2TeX is useful when larger chunks of text are more convenient to write in a Markdown-like syntax, while the target document should still remain a real LaTeX document. LaTeX commands, mathematical expressions, citations, and raw LaTeX environments are therefore intentionally allowed inside the Markdown.
A typical snippet:
# Motivation
The calibration follows @calibration and uses $\alpha < 0.1$ as threshold.
For the final selection we use:
- **nominal** reconstruction
- systematic variations with \cref{sec:systematics}
- the response model from [@response; @detector]
```tex
\begin{align}
y &= f(x) \\
&= x^2 + \alpha
\end{align}
```This turns into LaTeX output such as:
\section{Motivation}
The calibration follows \cite{calibration} and uses $\alpha < 0.1$ as threshold.
For the final selection we use:
\begin{itemize}
\item \textbf{nominal} reconstruction
\item systematic variations with \cref{sec:systematics}
\item the response model from \parencite{response, detector}
\end{itemize}
\begin{align}
y &= f(x) \\
&= x^2 + \alpha
\end{align}HTML comments remain in the generated .tex source as % comments and do
not appear in the PDF. Both single-line and multiline comments are supported:
<!-- This is a comment -->
<!--
TODO:
Complete this section later.
-->Comments can also occur within text. Each comment line receives a % prefix,
and following text resumes on a new source line so it stays visible. Markdown
and LaTeX inside comments are not parsed. An unclosed comment extends to the
end of the input. Comment syntax inside backtick code spans and fenced code
blocks stays literal.
The current parser intentionally supports only a small, tested subset:
- headings with
#,##,###, ... - horizontal rules with
---on its own line - paragraph detection for normal text
- HTML comments (
<!-- ... -->), preserved as LaTeX%comments - italics with
*text*or_text_ - bold with
**text**or__text__ - strikethrough with
~~text~~ - inline code with backticks, e.g.
`code` - fenced code blocks with three backticks
- fenced
texblocks as raw LaTeX output - blockquotes with
> - inline math with
$...$or\\(...\\) - display math with
$$...$$or\\[...\\] - pipe tables, including inline formatting and math inside cells
- LaTeX commands such as
\alpha,\cref{...}, or\textit{...} - unordered lists with
-,*, or+ - ordered lists with
1.or1) - some nested list forms
- single citations such as
@key - Pandoc-style citation groups such as
[@key1; @key2], including locators such as[@key, p. 433] - raw LaTeX environments with
\begin{...}and\end{...} - raw LaTeX environments with optional
$$ ... $$or\\[ ... \\]wrappers
The output is LaTeX-like:
- headings are mapped by default to
\section,\subsection,\subsubsection,\paragraph, and\subparagraph @keybecomes\cite{key}by default[@key1; @key2]becomes\parencite{key1, key2}by default[@key, p. 433]becomes\parencite[p. 433]{key}by default- lists become
itemizeorenumerate - blockquotes become LaTeX
quoteenvironments - pipe tables become full-width
tabularxenvironments with wrapping cells and automatic compact/text column selection; delimiter colons preserve alignment - normal code blocks are emitted as
verbatim texcode blocks are emitted unchanged as LaTeX$$...$$is normalized to\\[...\\];\\[...\\]is kept in that form- normal text is not automatically escaped for LaTeX special characters
Mathematics is a Mark2TeX extension; CommonMark itself does not define math delimiters. The supported forms and their output are:
| Input | Context | Output |
|---|---|---|
$x^2$ |
inline text, headings, lists, blockquotes, and table cells | unchanged |
\\(x^2\\) |
inline text, headings, lists, blockquotes, and table cells | unchanged |
$$x^2$$ |
display block, including lists and blockquotes | \\[x^2\\] |
\\[x^2\\] |
display block, including lists and blockquotes | unchanged |
\\begin{align}...\\end{align} |
raw LaTeX block | unchanged |
Markdown syntax inside math is not interpreted. LaTeX sub-environments that do
not open math mode themselves, such as aligned, cases, and matrix
environments, remain inside the normalized \\[...\\] display block. A
self-contained display environment such as align, gather, or equation is
emitted as that environment alone, avoiding an invalid nested display mode.
Display delimiters in pipe-table cells are kept literal and reported as
warnings; inline math is supported there.
Unclosed, empty, or mismatched math delimiters are preserved literally and
produce a math-delimiter warning during file conversion. Unclosed or
mismatched LaTeX environments produce a latex-environment warning. These
warnings include line and column when the position can be determined. A
delimiter preceded by an odd number of backslashes is treated as escaped. Code
spans, fenced code blocks, and tex blocks are never parsed as math.
A blockquote is written by starting every quoted line with >:
> **Beobachtungen** → Daten → erkennbare Zusammenhänge → Modell → Anwendung auf neue Fälle
>
> Eine zweite Zeile mit *Inline-Formatierung*.Mark2TeX removes the markers, parses the content as Markdown blocks, and
wraps the complete block in a LaTeX quote environment. This supports paragraphs,
tables, lists, fenced code, and nested blockquotes (> > ...), including the
usual inline formatting and math. Tables use the available width inside the quote.
Blank lines inside a blockquote must be written as > lines. Lazy continuation
lines without a > marker are not supported.
> **Tafel:** Zweispaltig sichern.
>
> | Daten | Modell |
> | --- | --- |
> | Eingabe $x$ | Parameter $w,b$ |
> | Zielwert $y$ | Vorhersage $\hat y=wx+b$ |Mark2TeX only transforms complete, unambiguous Markdown constructs. An
underscore in ordinary text (for example snake_case) is kept literally;
italics require a matching closing underscore or asterisk. Likewise, an
unclosed fenced code block or an unclosed \begin{...} environment is emitted
as ordinary text instead of being partially converted. Paragraphs immediately
before or after a list remain separate paragraphs.
Mark2TeX is not a full Markdown converter. In particular, do not expect the following to work like they do in CommonMark, Pandoc, or GitHub Markdown:
- links and images in Markdown syntax
- HTML blocks other than comments
- footnotes
- task lists
- reference links
- arbitrary escaping rules
- fully specified edge cases for nested inline elements
If you need such constructs, the preferred approach at the moment is usually
direct LaTeX, for example as a tex code block or a raw LaTeX environment.
In a LuaLaTeX document, Mark2TeX is loaded as a package:
\documentclass{article}
\usepackage{mark2tex}
\begin{document}
\mdinput{content.md}
\end{document}\mdinput{...} converts the given Markdown file to LaTeX and inputs the
generated file via \input. Generated files are placed in the mark2tex/
directory by default.
For chapter-like files there is also:
\mdinclude{chapter.md}This inputs the generated file via \include.
To compile only selected Markdown includes, use \mdincludeonly in the
preamble with the same Markdown paths passed to \mdinclude:
\mdincludeonly{introduction.md,conclusion.md}Like LaTeX's \includeonly, excluded files retain their auxiliary data. They
are not converted or registered as inputs during that run. An empty
\mdincludeonly{} excludes all \mdinclude files. \mdinput is unaffected.
Both \mdinput and \mdinclude register their Markdown source files with
LuaTeX's recorder. Build tools such as latexmk can therefore detect Markdown
changes from the generated .fls file and rebuild the document automatically.
Important: Mark2TeX requires LuaLaTeX. Other engines such as pdfLaTeX are not supported.
You can build a TeX Live package archive from this repository:
make distThis produces dist/mark2tex.tar.xz. The archive contains a TDS structure and
an embedded tlpobj, so it can be installed directly with tlmgr:
tlmgr install --file dist/mark2tex.tar.xzTo check the archive without installing it, use:
make tlmgr-install-dry-runThe package installs mark2tex.sty under tex/latex/mark2tex/ and the Lua
implementation under scripts/mark2tex/.
Overleaf usually does not allow project-local installation via tlmgr.
For Overleaf there is therefore a separate bundle:
make overleaf-zipThis produces dist/mark2tex-overleaf.zip. That archive is not a TeX Live
installation; instead it contains the files in a form that can be uploaded
directly into an Overleaf project:
mark2tex.sty
mark2tex.lua
src/mark2tex/*.lua
README.md
LICENSE
In Overleaf, mark2tex.sty and mark2tex.lua must live at the top level of the
project; the src/mark2tex/ directory must remain relative to them. In the
Overleaf menu, the compiler must be set to LuaLaTeX. After that, the package
can be used like it is locally:
\usepackage{mark2tex}
\begin{document}
\mdinput{content.md}
\end{document}Configure Mark2TeX directly when loading the package:
\usepackage[
citation=autocite,
paren-citation=parencite,
save-dir=generated-mark2tex,
header={chapter,section,subsection,subsubsection,paragraph}
]{mark2tex}| Option | Default | Meaning |
|---|---|---|
citation |
cite |
Command used for @key |
paren-citation |
parencite |
Command used for [@key] and citation groups |
save-dir |
mark2tex |
Directory for generated LaTeX files |
header |
{section,subsection,subsubsection,paragraph,subparagraph} |
Heading commands, from level 1 upward |
verbose |
false |
Log files when they are converted; use verbose or verbose=true to enable |
Command names are written without a leading backslash. Load any package that
provides your chosen citation commands separately. Enclose the comma-separated
header list in braces; it replaces the entire mapping. Deeper headings use
the last command in the list. Omitted or empty string options use their defaults.
The package no longer reads mark2tex_config.lua. Move existing settings into
package options, using save-dir for save_dir, paren-citation for
paren_citation, and a braced comma-separated list for header.
Inline LaTeX is preserved:
The corrected energy is $E_\mathrm{corr}$ and the result is shown in
\cref{fig:energy-response}.Citations can be written concisely:
The detector model follows @detector-note and the calibration strategy follows
[@calibration-paper; @run2-performance].For a parenthetical citation with a page or section locator, place the locator
after a comma. It is passed as the optional argument of \parencite:
The original proposal is discussed in [@Turing1950, p. 433].Pipe tables accept the usual alignment markers in their delimiter row and can contain the supported inline Markdown and LaTeX syntax:
| Quantity | Value | Comment |
| :------- | :---: | ------: |
| Energy | $E$ | **fit** |
| Events | 42 | @sample |This produces a tabularx spanning \linewidth, with left-, center-, and
right-aligned cells respectively. Font size stays unchanged. mark2tex.sty
automatically loads array, tabularx, and booktabs; standalone converter
output requires these packages in your document preamble.
Column selection is automatic and includes the header: a column is compact if its maximum cell length is at most 18 characters and its average at most 10. Markdown formatting does not count toward length; Unicode characters count once. Math and raw TeX use source length as a conservative approximation.
In mixed tables, compact columns use wrapping p{...} cells. Their widths are
computed from their longest cell and capped at half an equal column share of
the usable width (after intercolumn padding). Text columns share the remaining
space equally using X. If every column is compact, all columns use X, so
there is always a flexible column and no unused width. There are no weighted
text columns or per-table settings.
Tables use \toprule, \midrule, and \bottomrule, without outer column
padding or a center wrapper. Paragraph spacing separates tables from nearby
text. \linewidth also respects narrower containers such as minipages.
This handles ordinary prose by wrapping instead of scaling. Unbreakable words,
URLs, or long formulas can still overflow, and very many columns can become too
narrow. tabularx does not split tables across pages. For these cases, or for
captions and custom widths, use a raw LaTeX table.
More complex LaTeX blocks can be written directly in Markdown:
```tex
\begin{table}[h]
\centering
\caption{Nominal binning}
\label{tab:binning}
\begin{tabular}{l|c}
\toprule
Layer & bins \\
\midrule
1 & 32 \\
\bottomrule
\end{tabular}
\end{table}
```Raw environments without a code fence are also recognized:
\begin{align}
p(x) &= p(z)\left.\dv{f^{-1}}{x}\right|_{x=x_0}
\end{align}Project layout:
src/mark2tex/contains the Lua implementation.mark2tex.luais the compatibility entry point.mark2tex.styintegrates Mark2TeX into LuaLaTeX.tests/contains fixture tests, unit tests, and a LaTeX smoke test.scripts/contains development helpers.
Run regression tests:
make testRun individual fixtures:
lua tests/run.lua tests/header.test tests/lists.testTest LuaLaTeX integration:
make latex-smokeThe test runner uses luaunit if it is installed. If luaunit is missing, it
falls back to a small built-in assertion runner so that parser and writer
fixtures can still be checked.
At runtime, the project currently uses Lua modules for LPeg, filesystem access, and MD5:
lpeglfsmd5
A working LuaLaTeX installation is required for the LaTeX integration.
luaunit is optional for tests.