<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<oembed>
  <author_name>stockmark_tech</author_name>
  <author_url>https://blog.hatena.ne.jp/stockmark_tech/</author_url>
  <blog_title>Stockmark Tech Blog</blog_title>
  <blog_url>https://stockmark-tech.hatenablog.com/</blog_url>
  <categories>
  </categories>
  <description>ML事業部の近江崇宏です。 ストックマークではプロダクトで様々な自然言語処理の技術を用いていますが、その中のコア技術の一つに固有表現抽出があります。固有表現抽出はテキストの中から固有表現（固有名詞）を抽出する技術で、例えば「Astrategy」というプロダクトでは、固有表現抽出を用いてニュース記事の中から企業名を抽出しています。（企業名抽出については過去のブログ記事を参考にしてください。） 一般に、固有表現抽出を行うためには、大量のテキストに固有表現をアノテーションした学習データをもとに機械学習モデルの学習を行います。今回、ストックマークは固有表現抽出のための日本語の学習データセットを公開いた…</description>
  <height>190</height>
  <html>&lt;iframe src=&quot;https://hatenablog-parts.com/embed?url=https%3A%2F%2Fstockmark-tech.hatenablog.com%2Fentry%2F2020%2F12%2F15%2F120000&quot; title=&quot;Wikipediaを用いた日本語の固有表現抽出データセットの公開 - Stockmark Tech Blog&quot; class=&quot;embed-card embed-blogcard&quot; scrolling=&quot;no&quot; frameborder=&quot;0&quot; style=&quot;display: block; width: 100%; height: 190px; max-width: 500px; margin: 10px 0px;&quot;&gt;&lt;/iframe&gt;</html>
  <image_url>https://cdn-ak.f.st-hatena.com/images/fotolife/s/stockmark_tech/20240416/20240416184132.jpg</image_url>
  <provider_name>Hatena Blog</provider_name>
  <provider_url>https://hatena.blog</provider_url>
  <published>2020-12-15 12:00:00</published>
  <title>Wikipediaを用いた日本語の固有表現抽出データセットの公開</title>
  <type>rich</type>
  <url>https://stockmark-tech.hatenablog.com/entry/2020/12/15/120000</url>
  <version>1.0</version>
  <width>100%</width>
</oembed>
