Monkt logo
Document Processing Files · Ocr Doc Parsing

Monkt

Monkt converts PDF, Word, Excel, PowerPoint, CSV, HTML files and web page URLs into clean Markdown or structured JSON designed for large language models, through a web dashboard or an authenticated REST API.

Active GDPR compliant Free trial Subscription API available Verified by Guidaio
Overview

What is Monkt?

Monkt is a document processing platform aimed at organisations and developers who feed content to AI systems. Its job is narrow on purpose: take a document in a format people actually work with, and return either clean Markdown or structured JSON that a large language model can consume without further cleanup.

The input list covers PDF, Word, PowerPoint, Excel, CSV, raw HTML, live web pages fetched from a URL, and images. A conversion hub exposes twelve named routes, six towards Markdown and six towards JSON, so every pairing has its own entry point. JSON output relies either on automatic schema detection or on a schema the customer defines, which is what makes the result predictable enough to drop into a pipeline.

DeepExtract is the premium tier of that engine. It runs several specialised models against what defeats plain text extraction: tables, figures, mathematical formulas, code blocks, citations and scanned pages. It returns a ZIP archive holding high-quality page images, extracted tables, page screenshots and additional formats, with images delivered either base64-embedded or as separate referenced files. A boundary detection system maps the coordinates of every table and image, OCR covers scanned material, and a deterministic JSON export comes in two shapes: a streamlined text version, and a structured one preserving the full document hierarchy. Page images are presented as compatible with the ColPali and LitePali libraries. DeepExtract is open to Pro and Enterprise subscribers, the latter on GPU-accelerated servers.

Access runs two ways. The dashboard offers drag-and-drop upload with a real-time preview panel. The REST API authenticates with a token in the Authorization header and exposes endpoints to list, create and delete transformations, with pagination, field filtering and ordering. Five documented recipes show the intended uses: invoices to JSON, articles to JSON, research papers to JSON, document processing inside AI agents, and preparation for LLM fine-tuning. A direct CrewAI integration, an SDK and webhook callbacks are cited for agent workflows.

The homepage claims more than one million documents processed, over ten supported formats and more than seven thousand users. The publisher is Premium Software Ltd., based in Sofia, Bulgaria, and founded by Simeon Emanuilov, PhD.

What it does

  • Convert PDF, Word, PowerPoint, Excel, CSV and HTML files into clean Markdown
  • Turn the same documents into structured JSON, using automatic schema detection or a custom schema
  • Transform a live web page into Markdown or JSON straight from its URL
  • Extract images embedded in documents and describe their content as text or structured data
  • Run OCR over scanned documents, from the Pro plan onwards
  • Process batches of documents in parallel, with progress tracking and notifications
  • Drive all of it programmatically through a token-authenticated REST API
Audience

When to use Monkt / When not to

A quick filter to help you decide if Monkt is the right fit.

When to use Monkt

  • Engineering teams building RAG pipelines who need consistent, chunk-ready text out of messy source documents
  • Developers automating document ingestion through a REST API rather than a user interface
  • Academic researchers turning papers into structured JSON with sections, references and figures
  • Finance and accounting teams extracting fields from invoices into a predictable schema
  • Knowledge workers and Obsidian users importing PDFs and web pages into a personal knowledge base

When not to use Monkt

  • Anyone wanting a permanently free tool: the cheapest tier costs USD 4.99 per month and no free plan exists
  • Mobile-first users, since Monkt ships no iOS or Android application
  • Teams needing audio or video transcription, which the site lists only as coming soon
  • Buyers with strict procurement requirements, as no DPA, subprocessor list or hosting location is published
  • People looking for a writing or editing assistant: Monkt converts documents, it never authors them
Get started

How to use Monkt

A typical end-to-end flow, from setup to results.

  1. Drop up to three files of 5 MB each on the homepage, or paste one or more URLs, then press Transform now
  2. Create an account to store your documents and retrieve them later
  3. Pick the output format: Markdown, or JSON with automatic schema detection
  4. On a Pro plan, supply your own JSON schema when you need the output shaped a specific way
  5. Follow the conversion in the real-time preview panel of the dashboard
  6. For DeepExtract jobs, download the ZIP archive holding extracted images, tables, page screenshots and the extra formats
  7. To automate, collect your API token and send it as an Authorization: Token header
  8. POST a file to the transformations endpoint, or POST a JSON body carrying a url field to convert a web page
  9. List, inspect or delete past transformations through the matching GET and DELETE endpoints
  10. Page, filter and sort large result sets with the page, page_size, field filter and ordering parameters
Quick read

Pros & Cons

Pros

  • Broad format coverage for a tool of this size, with twelve named conversion routes
  • Two complementary outputs: readable Markdown and JSON whose shape you control
  • Publicly documented REST API with curl examples, readable without an account
  • Low entry price at USD 4.99 per month
  • Smart caching means repeated conversions are served from cache without consuming the monthly quota
  • A European publisher with legal name, registration number and postal address in the open
  • Explicit commitment not to train models on customer data, and immediate deletion of the original files

Cons

  • No permanent free plan, and a tight entry quota of 50 transformations per month with a 15 MB file limit
  • Full API access and custom JSON schemas are reserved for the Pro and Enterprise plans
  • Retention is stated three different ways across the site: 7 or 30 days by plan, 30 days for everyone, and 24 hours
  • A SOC 2 claim appears in a single FAQ, with no trust page, no report and no subprocessor list
  • No data processing agreement is published or mentioned, and no hosting country or region is disclosed
  • Refund terms are unusually firm: any started period is due and unused credits are lost
  • Support carries no response-time commitment outside the Enterprise plan
Pricing

Pricing & Plans

Monkt does not offer a permanent free plan. The service can be tried without a subscription on up to three files of 5 MB each from the homepage, and the FAQ confirms that a sample document may be converted before committing. The lowest paid entry point is the Start plan at USD 4.99 per month. Pro is priced at USD 14.99 per month, while the Enterprise and Custom tiers are quoted on request.

Start, USD 4.99 per month, presented for individuals and small projects
  • 50 transformations per month
  • files up to 15 MB
  • 7 days of data persistence
  • export to Markdown and JSON
  • document storage and retrieval
Enterprise, priced on request
  • unlimited data persistence
  • faster inference on GPU servers
  • advanced DeepExtract processing
  • page layout
  • reading order and table structure understanding
  • custom integration support
  • extra long context window
Custom, quoted per request and badged Premium
  • bulk document processing and cleaning
  • intelligent chunking and vectorisation
  • RAG-optimised data preparation
  • custom chatbot data pipelines
  • dedicated project management
  • university and enterprise discounts
  • secure data handling and compliance
Special offers — University and enterprise discounts are listed among the Custom plan features, with no published rate, amount or eligibility condition
Prices and plans listed above may evolve. Always check the official pricing page before subscribing.
Trust & Privacy

Data, GDPR & hosting

A consolidated view of how Monkt handles your data.

GDPR overview

Monkt is published by a company established inside the European Union, in Sofia, Bulgaria, so no Article 27 representative is required and none is named. A dedicated privacy address, privacy@monkt.com, is published, and the privacy policy, last updated on 15 April 2026, lists rights of deletion, access, export and account closure. Concrete measures are stated: encryption in transit and at rest, isolated processing environments, access controls and authentication, plus regular security monitoring. A recipe page claims GDPR-compliant data handling. The gaps are real, however: no data processing agreement is offered or even mentioned, no subprocessor list is published, no data protection officer is named, and no hosting jurisdiction is disclosed. A SOC 2 claim appears in one FAQ without any supporting report.

Who owns the data?

The terms grant the user a limited, non-exclusive, non-transferable licence to use Monkt for personal or business purposes, and claim no ownership over the documents submitted or the converted output. The obligation runs the other way: the user warrants holding the right to process the files uploaded, and remains responsible for keeping their own backups. On the vendor side, the privacy policy states that original files are deleted immediately after transformation, that only the converted output is retained for the duration attached to the plan, and that customer data is not used to train models. Permanent deletion is available on request, alongside rights of access, export and account closure.

Reuse rights

Nothing in the terms restricts what the user may do with the converted output: Markdown and JSON exports are available from the entry-level Start plan onwards, and the licence explicitly covers business as well as personal use, so the results can be reused without asking permission. The restrictions imposed run in the opposite direction and concern the service itself: no reverse engineering, no breach of the file processing limits, no illegal use and no infringement of third-party intellectual property. Two practical limits are worth noting, both of them commercial rather than legal: full API access is reserved for the Pro and Enterprise plans, and custom JSON schemas are a Pro feature.

Data retention & training

Retention summary
Original files are never kept: the privacy policy states they are processed and then deleted immediately after transformation. Only the converted output is stored, for a period tied to the plan, 7 days on Start, 30 days on Pro, and a contractually agreed duration on Enterprise. The pricing table matches that split and advertises unlimited persistence on Enterprise. Cached results follow the same window as the plan. Users can request permanent deletion at any time, and may access, export or close their account. Two caveats: the homepage FAQ instead states that documents are deleted after 30 days unless specified otherwise, and one recipe page claims all data is purged after 24 hours, so the published durations contradict each other. No anonymisation is mentioned anywhere.
Trains on customer data
No
GDPR contact

Hosting summary

Monkt does not disclose where user data is hosted. No country, region, data centre or infrastructure provider is named anywhere on the site, and there is no trust or security page. What is known is the publisher's own jurisdiction: Premium Software Ltd. is established in Sofia, Bulgaria, inside the European Union. Stripe is named as the payment processor in both the privacy policy and the terms, and is the only third party identified. The security measures stated are encryption in transit and at rest, secure and isolated processing environments, access controls with authentication, and regular security updates and monitoring. One structural point limits exposure: original files are not hosted for any length of time, since they are deleted immediately after transformation, and only the converted output is stored. Enterprise customers are told private deployments with additional security measures are possible. Note that the domain resolves behind a Cloudflare anycast address, which says nothing about where the data itself actually rests.

Watch-outs

Things to keep in mind

Risks and trade-offs to weigh before adopting Monkt.

  • Confidential material leaves your control: contracts, invoices and research files all transit through a third-party service
  • Retention is announced three incompatible ways, so you cannot actually know when your converted output disappears
  • The SOC 2 mention rests on a single FAQ line with no report or auditor named, and should not be read as a verified audit
  • With no data processing agreement available, a GDPR controller lacks the contract required to use a processor lawfully
  • Automated extraction is fallible: invoice fields, tables and formulas need a human check before any accounting or legal use
  • Building an ingestion pipeline on a single external service creates one point of failure you do not operate
  • The terms are unforgiving on money: any started billing period is due, unused credits are lost, and a chargeback raised before contacting support can suspend the account
Setup

Setup & Integrations

Technical difficulty

Basic use requires no technical skill: drag a file onto the homepage or paste a URL, press Transform now, and read the result. An account is only needed to store and retrieve documents. The API raises the bar modestly, to the level of anyone who has called a REST endpoint: collect a token, send it as an Authorization header, and follow the curl examples in the public documentation. Writing a custom JSON schema is the one step assuming real familiarity with structured data, and automatic detection exists for those who would rather skip it. Nothing is installed.

Deployment

Web appAPI

Integrations

CrewAI Obsidian
Company

Behind Monkt

Company name
Premium Software Ltd.
Founded
05/08/2018
Country of origin
🇧🇬 Bulgaria
Headquarters
str. Vasil Petleshkov 80, Sofia, Bulgaria
UBO
Simeon Emanuilov
UBO country
🇧🇬 Bulgaria
Domain registrar country
🇺🇸 United States
Legal contact
Support contact

Social

Official links

Resources

All the official URLs gathered for verification and reference.

FAQ

Frequently asked questions

Which file formats can Monkt process?
PDF, Word in DOC and DOCX, PowerPoint in PPT and PPTX, Excel in XLS and XLSX, CSV, HTML and plain text, plus images in JPG, PNG and WEBP. A web page can also be converted directly from its URL. Text and embedded images are both handled.
Is there a free plan?
No. The pricing table lists four paid tiers and the cheapest, Start, costs USD 4.99 per month. You can still test the service before subscribing: the homepage accepts up to three files of 5 MB each, and the FAQ confirms a sample document can be converted first.
How does custom JSON schema output work?
Pro users define a schema that states exactly how the data should be structured, or rely on automatic schema detection instead. The point of the custom route is a predictable output shape that downstream code can depend on.
Does Monkt have an API?
Yes. A REST API authenticates with a token in the Authorization header and exposes endpoints to create, list, inspect and delete transformations, with pagination, field filtering and ordering. Full API access is reserved for the Pro and Enterprise plans.
How long is my data kept?
The privacy policy states that original files are deleted immediately after transformation, and that only the converted output is retained: 7 days on Start, 30 days on Pro, and a contractual duration on Enterprise. Be aware that the homepage FAQ instead announces 30 days for everyone, and one recipe page mentions a 24-hour purge, so the three statements do not agree.
Are my documents used to train AI models?
No. The privacy policy states plainly that Monkt does not train its models on customer data. No opt-out setting is documented, because on that stated position there is nothing to opt out of.
What is DeepExtract?
It is the premium extraction tier, running several specialised models over tables, figures, mathematical formulas, code blocks, citations and scanned pages. Results arrive as a ZIP archive containing page images, extracted tables, screenshots and additional formats, including a deterministic JSON export. It is available on the Pro and Enterprise plans.
What support is included?
All users get documentation and email support, Pro adds priority support, and Enterprise customers get dedicated teams and custom SLAs. The terms are more restrictive than the marketing: outside Enterprise, support is provided on a best-availability basis with no guaranteed response time.
Who is behind Monkt?
Premium Software Ltd., registration number BG207947833, based at str. Vasil Petleshkov 80 in Sofia, Bulgaria. The platform was founded by Simeon Emanuilov, PhD and senior software engineer.
Is there a mobile app?
No. Monkt runs as a web application and a REST API, and the site links to no iOS or Android application. A VS Code extension and an Obsidian plugin appear on the 2026 roadmap but are not shipped.
Conclusion

Should you pick Monkt?

Monkt does one thing and states it plainly: it turns documents people actually work with into formats a language model can read. The format coverage is broad for a product this size, the dual Markdown and JSON output is genuinely useful, the REST API is documented in the open with working examples, and the entry price of USD 4.99 per month puts it within reach of an individual. DeepExtract is the part that separates it from a simple converter, with table and figure boundary detection, OCR and a deterministic JSON export intended for pipelines that cannot tolerate drift.

The reservations are as clear as the strengths. There is no permanent free plan, the entry quota is tight, and the two features most likely to matter to a technical buyer, full API access and custom JSON schemas, sit behind the Pro tier. More importantly, the trust story is unfinished: data retention is described three different ways across three pages, a SOC 2 claim appears once in a recipe FAQ with nothing to back it, no data processing agreement exists, and no hosting country is ever named. Against that, the commitment not to train on customer data and the immediate deletion of original files are real and clearly stated, and the publisher is identifiable: a Bulgarian company with a registration number, a postal address and a named founder.

Treat Monkt as a capable specialist tool rather than a compliance-ready enterprise platform. For RAG teams, developers, researchers and document-heavy professionals who can verify the output, it earns its place. Anyone facing a procurement questionnaire should ask for the retention rules and a DPA in writing first.