개발자를 위한 기술 SEO 크롤링 및 인덱스 체크리스트

작성자

카테고리:

← 피드로
DEV Community · rokya elbarbary · 2026-09-29 개발(SW)

rokya elbarbary

Search visibility starts with three separate questions, and a page can fail any one of them independently:

  1. Can a search engine discover and fetch the URL? (crawling)
  2. Can it see the content once fetched? (rendering)
  3. Does it choose to keep the page in its index, and under which URL? (indexing and canonicalisation)

Work through the checklist in that order. There is little point tuning titles on a page that is blocked in robots.txt.

1. Crawling

Check How to verify What good looks like robots.txt is reachable Request /robots.txt directly Returns HTTP 200 (or 404 if you intentionally have none); never a 5xx, which can make Google pause crawling Important paths are not disallowed Read every Disallow rule; test key URLs in Search Console’s robots.txt report Templates, CSS and JS needed for rendering are crawlable XML sitemap exists and is referenced Sitemap: line in robots.txt; submitted in Search Console Lists only canonical, indexable URLs that return 200 Internal links use real anchors Inspect navigation and pagination <a href="..."> elements, not click handlers on <div> or <span> No redirect chains on internal links Crawl the site with a crawler of your choice Internal links point straight at the final URL Server responds consistently Check server logs for Googlebot requests Stable 200 responses; no bursts of 429 or 5xx

2. Rendering

Modern sites often build content in the browser. Google can render JavaScript, but rendering happens after the initial fetch and depends on every resource loading successfully.

  • Compare view-source with the rendered DOM (URL Inspection tool, “View crawled page”). Primary content, headings and internal links should exist in the rendered HTML.
  • Do not block the JavaScript or CSS files your templates depend on.
  • Avoid content that only appears after a user interaction (clicking a tab, scrolling an infinite list) if you need it indexed. Put it in the HTML or render it server-side.
  • Lazy-loaded images should use native loading="lazy" or a method that exposes the image URL in the rendered DOM.

3. Indexing and canonicalisation

Check What to look for meta robots and X-Robots-Tag No accidental noindex left over from staging; check both the HTML and the HTTP headers Canonical tags One rel="canonical" per page, absolute URL, pointing at a 200, indexable URL Duplicate variants http vs https, www vs non-www, trailing slash, uppercase, tracking parameters – each should resolve to one canonical version via redirect or canonical Status codes Removed pages return 404 or 410; moved pages return 301 to the closest equivalent Soft 404s Thin “no results” or empty category pages returning 200 – either add content or return 404 Language versions If you publish Arabic and English versions, use hreflang annotations that are reciprocal and point at canonical URLs

4. Page-level basics

  • A unique, descriptive <title> and meta description on every indexable template. Placeholder descriptions left over from a CMS theme are a common and easily fixed problem.
  • Exactly one visible H1 that describes the page. Duplicate or placeholder headings (for example a leftover “Heading” block) confuse users and are worth removing.
  • Descriptive internal anchor text rather than “click here”.

5. After launch

  1. Submit the sitemap and inspect a sample of key URLs in Search Console.
  2. Watch the Pages (indexing) report for spikes in “Crawled – currently not indexed”, “Duplicate without user-selected canonical” or “Blocked by robots.txt”.
  3. Re-run the crawl after every major release. Technical SEO regressions usually arrive with deployments, not with content.

Further reading

원문에서 계속 ↗